3rd Year Undergraduate Researcher at IISER Bhopal

Building AI that
keeps listening.

I study how audio, language, and vision can work together when sound is incomplete, noisy, or constantly changing.

01Featured research

Accepted · CIKM 2026 · Rome

PRISM

Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

A gradient-free test-time adaptation framework for improving audio-text representations when the acoustic world changes.

Read the research

02 · Research

Research
areas.

01

Multimodal AI

Learning representations across audio, vision, and language for multimodal understanding.

02

Speech & audio

Studying speech representations, prosody, and audio-language systems under changing acoustic conditions.

03

Efficient sequence modeling

Exploring state-space models and efficient architectures for long-sequence and multimodal learning.

03 · Research & Projects

Research &
projects.

Speech research2025 — 2026

Pitch Stylization

Information-theoretic pitch stylization and prosodic feature extraction using dynamic programming and statistical model selection.

ProsodyInformation theory
Research project01

Mamba-AVSR

Audio-Visual Speech Recognition with State Space Models. An experimental pipeline combining Whisper-v3, VideoMAE, Mamba-2, and Llama-3.2.

PyTorchAVSRMamba-2
Research project02

VaaniAdapt

A rehearsal-free, parameter-efficient continual learning framework for extending a deployed multilingual ASR system to new languages without storing prior speech data.

Continual learningLow-rank adaptersMultilingual ASRParameter-efficient
Applied AI03

DermaScribe

AI-powered skincare analysis and recommendation platform with computer vision, SQLite + FTS5, caching, and Streamlit.

04 · About

About
Ashish.

I am a Data Science and Engineering undergraduate at the Indian Institute of Science Education and Research, Bhopal. I build AI systems for speech, audio, and multimodal understanding.

Recent work includes information-theoretic pitch stylization at LTRC, IIIT Hyderabad, ASR continual learning, and audio-visual sequence modeling. I am currently a remote research intern with the NTU Speech Lab.

Multimodal Machine LearningTest-time adaptationInformation theorySpeech & Audio ProcessingEfficient AIState-space models

05 · Experience

Where the work
has taken me.

2026

Language Technologies Research Centre

Research Intern · IIIT Hyderabad

2026

NTU Speech Lab

Research Intern · Remote · Ongoing

2025–26

VisDom

Research Intern

2024–28

IISER Bhopal

B.Tech in Data Science and Engineering

06 · Contact

Get in
touch.

For research discussions, collaborations, or internship opportunities.