Research
Selected publications by theme, newest first. * marks equal contribution. The full list is on Google Scholar.
Audio understanding and agents
- HARP: agentic hybrid retrieval and analysis for long-form audio Chin-Jou Li, Masao Someki, Woojeong Jin, Yashish M. Siriwardena, Tanmay Laud, Shanil Puri, Shinji Watanabe. arXiv preprint arXiv:2609.14116. [Preprint]
Speech inversion and production
- Speaker-independent speech inversion for recovery of velopharyngeal port constriction degree Yashish M. Siriwardena, Suzanne E. Boyce, Mark K. Tiede, Liran Oren, Brittany Fletcher, Michael Stern, Carol Y. Espy-Wilson. Journal of the Acoustical Society of America 156(2): 1380-1390.
- Improving speech inversion through self-supervised embeddings and enhanced tract variables Yashish M. Siriwardena*, Ahmed Adel Attia*, Carol Espy-Wilson. European Signal Processing Conference (EUSIPCO) 2024.
- Speaker-independent speech inversion for estimation of nasalance Yashish M. Siriwardena, Carol Espy-Wilson, Suzanne Boyce, Mark K. Tiede, Liran Oren. Interspeech 2023. [Preprint]
- The Secret Source: incorporating source features to improve acoustic-to-articulatory speech inversion Yashish M. Siriwardena, Carol Espy-Wilson. IEEE ICASSP 2023. [Paper][Poster][Video][Code]
- Audio data augmentation for acoustic-to-articulatory speech inversion Yashish M. Siriwardena, Ahmed Adel Attia, Ganesh Sivaraman, Carol Espy-Wilson. European Signal Processing Conference (EUSIPCO) 2023. [Preprint]
- Acoustic-to-articulatory speech inversion with multi-task learning Yashish M. Siriwardena, Ganesh Sivaraman, Carol Espy-Wilson. Interspeech 2022. [Paper][Slides]
Synthesis and voice conversion
- Accent conversion with articulatory representations Yashish M. Siriwardena, Nathan Swedlow, Audrey Howard, Evan Gitterman, Dan Darcy, Carol Espy-Wilson, Andrea Fanelli. Interspeech 2024.
- Learning to compute the articulatory representations of speech with the MirrorNet Yashish M. Siriwardena, Carol Espy-Wilson, Shihab Shamma. Interspeech 2023. [Preprint]
- The MirrorNet: learning audio synthesizer controls inspired by sensorimotor interactions Yashish M. Siriwardena, Guilhem Marion, Shihab Shamma. IEEE ICASSP 2022. [Paper][Poster][Video][Code]
Speech and mental health
- A multimodal framework for assessment of the schizophrenia spectrum Gowtham Premananth, Yashish M. Siriwardena, Philip Resnik, Sonia Bansal, Deanna L. Kelly, Carol Espy-Wilson. Interspeech 2024.
- A multi-modal approach for identifying schizophrenia using cross-modal attention Gowtham Premananth, Yashish M. Siriwardena, Philip Resnik, Carol Espy-Wilson. IEEE Engineering in Medicine and Biology Society (EMBC) 2024.
- Acoustic-to-articulatory speech inversion features for mispronunciation detection of /ɹ/ in child speech sound disorders Yashish M. Siriwardena*, Nina R. Benway*, Jonathan L. Preston, Elaine Hitchcock, Tara McAllister, Carol Espy-Wilson. Interspeech 2023. [Preprint]
- Multimodal approach for assessing neuromotor coordination in schizophrenia using convolutional neural networks Yashish M. Siriwardena, Carol Espy-Wilson, Deanna L. Kelly, Chris Kitchen. ACM International Conference on Multimodal Interaction (ICMI) 2021. [Paper][Poster][Video]
- Emotion recognition with articulatory coordination features Yashish M. Siriwardena. 181st Meeting of the Acoustical Society of America 2021. [Poster]
- Inverted vocal tract variables and facial action units to quantify neuromotor coordination in schizophrenia Yashish M. Siriwardena. 12th International Seminar on Speech Production (ISSP) 2020. [Preprint]