Multimodal AI
Learning representations across audio, vision, and language for multimodal understanding.
3rd Year Undergraduate Researcher at IISER Bhopal
I study how audio, language, and vision can work together when sound is incomplete, noisy, or constantly changing.
Accepted · CIKM 2026 · Rome
Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift
A gradient-free test-time adaptation framework for improving audio-text representations when the acoustic world changes.
Read the research ↗02 · Research
Learning representations across audio, vision, and language for multimodal understanding.
Studying speech representations, prosody, and audio-language systems under changing acoustic conditions.
Exploring state-space models and efficient architectures for long-sequence and multimodal learning.
03 · Research & Projects
Information-theoretic pitch stylization and prosodic feature extraction using dynamic programming and statistical model selection.
Audio-Visual Speech Recognition with State Space Models. An experimental pipeline combining Whisper-v3, VideoMAE, Mamba-2, and Llama-3.2.
A rehearsal-free, parameter-efficient continual learning framework for extending a deployed multilingual ASR system to new languages without storing prior speech data.
AI-powered skincare analysis and recommendation platform with computer vision, SQLite + FTS5, caching, and Streamlit.
04 · About
I am a Data Science and Engineering undergraduate at the Indian Institute of Science Education and Research, Bhopal. I build AI systems for speech, audio, and multimodal understanding.
Recent work includes information-theoretic pitch stylization at LTRC, IIIT Hyderabad, ASR continual learning, and audio-visual sequence modeling. I am currently a remote research intern with the NTU Speech Lab.
05 · Experience
Research Intern · IIIT Hyderabad
Research Intern · Remote · Ongoing
Research Intern
B.Tech in Data Science and Engineering
06 · Contact
For research discussions, collaborations, or internship opportunities.