English
Related papers

Related papers: A Dataset for measuring reading levels in India at…

200 papers

Many commercial and forensic applications of speech demand the extraction of information about the speaker characteristics, which falls into the broad category of speaker profiling. The speaker characteristics needed for profiling include…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-14 Shareef Babu Kalluri , Deepu Vijayasenan , Sriram Ganapathy , Ragesh Rajan M , Prashant Krishnan

Children's speech recognition is considered a low-resource task mainly due to the lack of publicly available data. There are several reasons for such data scarcity, including expensive data collection and annotation processes, and data…

Computation and Language · Computer Science 2024-06-25 Vrunda N. Sukhadia , Shammur Absar Chowdhury

Streaming end-to-end automatic speech recognition (ASR) models are widely used on smart speakers and on-device applications. Since these models are expected to transcribe speech with minimal latency, they are constrained to be causal with…

Recent advancements in text-to-speech (TTS) synthesis show that large-scale models trained with extensive web data produce highly natural-sounding output. However, such data is scarce for Indian languages due to the lack of high-quality,…

Bengali is spoken by over 230 million people yet remains severely under-served in automatic speech recognition (ASR) and speaker diarization research. In this paper, we present our system for the DL Sprint 4.0 Bengali Long-Form Speech…

Computation and Language · Computer Science 2026-03-23 Md. Nazmus Sakib , Shafiul Tanvir , Mesbah Uddin Ahamed , H. M. Aktaruzzaman Mukdho

Early detection of asthma in children is crucial to prevent long-term respiratory complications and reduce emergency interventions. This work presents an AI-powered diagnostic pipeline that leverages Googles Health Acoustic Representations…

Sound · Computer Science 2025-04-30 Abul Ehtesham , Saket Kumar , Aditi Singh , Tala Talaei Khoei

Abusive content detection in spoken text can be addressed by performing Automatic Speech Recognition (ASR) and leveraging advancements in natural language processing. However, ASR models introduce latency and often perform sub-optimally for…

Sound · Computer Science 2022-02-17 Vikram Gupta , Rini Sharon , Ramit Sawhney , Debdoot Mukherjee

This paper describes the systems developed by SPRING Lab, Indian Institute of Technology Madras, for the ASRU MADASR 2.0 challenge. The systems developed focuses on adapting ASR systems to improve in predicting the language and dialect of…

Computation and Language · Computer Science 2025-11-20 Arjun Gangwar , Kaousheik Jayakumar , S. Umesh

Automatic speech recognition (ASR) has been significantly advanced with the use of deep learning and big data. However improving robustness, including achieving equally good performance on diverse speakers and accents, is still a…

Sound · Computer Science 2020-11-17 Fan Yu , Zhuoyuan Yao , Xiong Wang , Keyu An , Lei Xie , Zhijian Ou , Bo Liu , Xiulin Li , Guanqiong Miao

Automatic Speech Recognition (ASR) systems are known to exhibit difficulties when transcribing children's speech. This can mainly be attributed to the absence of large children's speech corpora to train robust ASR models and the resulting…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-22 Jenthe Thienpondt , Kris Demuynck

The performance of automated speech recognition (ASR) systems is well known to differ for varied application domains. At the same time, vendors and research groups typically report ASR quality results either for limited use simplistic…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-31 Rostislav Kolobov , Olga Okhapkina , Olga Omelchishina , Andrey Platunov , Roman Bedyakin , Vyacheslav Moshkin , Dmitry Menshikov , Nikolay Mikhaylovskiy

Automatic Speech Recognition (ASR) has increasing utility in the modern world. There are a many ASR models available for languages with large amounts of training data like English. However, low-resource languages are poorly represented. In…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-24 Kavitha Raju , Anjaly V , Ryan Lish , Joel Mathew

Researchers have recently started to study how the emotional speech heard by young infants can affect their developmental outcomes. As a part of this research, hundreds of hours of daylong recordings from preterm infants' audio environments…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-18 Einari Vaaras , Sari Ahlqvist-Björkroth , Konstantinos Drossos , Okko Räsänen

We developed dysarthric speech intelligibility classifiers on 551,176 disordered speech samples contributed by a diverse set of 468 speakers, with a range of self-reported speaking disorders and rated for their overall intelligibility on a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-17 Subhashini Venugopalan , Jimmy Tobin , Samuel J. Yang , Katie Seaver , Richard J. N. Cave , Pan-Pan Jiang , Neil Zeghidour , Rus Heywood , Jordan Green , Michael P. Brenner

This study investigates the performance of personalized automatic speech recognition (ASR) for recognizing disordered speech using small amounts of per-speaker adaptation data. We trained personalized models for 195 individuals with…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-12 Jimmy Tobin , Katrin Tomanek

Modern Automatic Speech Recognition (ASR) technology has evolved to identify the speech spoken by native speakers of a language very well. However, identification of the speech spoken by non-native speakers continues to be a major challenge…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Afroz Ahamad , Ankit Anand , Pranesh Bhargava

Sign languages are the primary means of communication for many hard-of-hearing people worldwide. Recently, to bridge the communication gap between the hard-of-hearing community and the rest of the population, several sign language…

Computation and Language · Computer Science 2023-07-12 Abhinav Joshi , Susmit Agrawal , Ashutosh Modi

An independent, automated method of decoding and transcribing oral speech is known as automatic speech recognition (ASR). A typical ASR system extracts feature from audio recordings or streams and run one or more algorithms to map the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-21 Tushar Talukder Showrav

Reading fluency assessment is a critical component of literacy programmes, serving to guide and monitor early education interventions. Given the resource intensive nature of the exercise when conducted by teachers, the development of…

Computation and Language · Computer Science 2024-06-04 Mithilesh Vaidya , Binaya Kumar Sahoo , Preeti Rao

Automatic reading aloud evaluation can provide valuable support to teachers by enabling more efficient scoring of reading exercises. However, research on reading evaluation systems and applications remains limited. We present a novel…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-01 Lingyun Gao , Cristian Tejedor-Garcia , Catia Cucchiarini , Helmer Strik