Kaveri K. Sheth
/ˈkɑːvəɾi ʃɛθ/ · Paris, France
Computational linguist & developmental language researcher. I study how infants learn language from the sparse, noisy, multimodal input of everyday life and build the speech pipelines, benchmarks, and open infrastructure that make that science possible. Postdoctoral researcher at LSCP — ENS, EHESS, CNRS, PSL University, Paris.
About
My work sits at the intersection of AI and the science of child language learning. I use automatic speech recognition and data science to characterize the ecological input children actually learn from — building speech-processing pipelines and machine learning benchmarks from long-form, naturalistic audio recordings across typologically diverse languages, and validating them cross-linguistically.
I am currently a postdoctoral researcher working with Dr. Alejandrina Cristia at Ecole Normale Superieure. I hold a PhD in Linguistics and an MS in Computational Linguistics from the University of Washington, where I worked with Drs. Naja Ferjan Ramírez, Patricia Kuhl, and Gina-Anne Levow. I publish in both machine learning venues (Interspeech) and cognitive science journals, and I lead ELSI, a nonprofit sustaining open research infrastructure for the language development community.
Beyond the research itself, I care deeply about how academia trains its people. I've mentored 15+ students, and I'm currently writing a book on mentorship in academia — what it actually means to be a good mentor, and how we could do it better. I write about this and more on my Substack.
Publications
Full list on Google Scholar →
Projects & leadership
ELSI
Nonprofit sustaining multimodal, open research infrastructure for the language development community. As CEO, I lead funding, hiring, and partnerships with research teams worldwide.
nonprofit · open science · ceoChild-directed speech detection
Rule-based system for separating child-directed from overheard speech in long-form recordings, validated across typologically diverse datasets with collaborators on four continents.
python · asr · cross-linguisticLong-form benchmarking framework
A framework for deriving machine learning benchmarks from naturalistic recordings — enabling principled comparison of how humans and machines learn from similar input.
benchmarks · interspeech 2026Gendered citation tracking
Python NLP pipeline tracking gendered citation patterns in news media at scale, built during a data science internship at ReThink Media.
spacy · neuralcoref · pandasTalks, posters & service
Teaching & mentorship
Honors
Writing
A book on mentorship in academia
In progress. What it actually means to be a good mentor — and how academia, which runs on mentorship but rarely teaches it, could do better by the people it trains.
book · in progressSubstack
Essays on mentorship, academia, language science, and life as a researcher. New posts land in your inbox — subscribe to follow along.
substack · essaysContact
I'm currently exploring [faculty / research scientist / industry research] opportunities. If you're hiring — or want to talk about child language, speech technology, or open research infrastructure — my inbox is open.