English
Related papers

Related papers: Vakyansh: ASR Toolkit for Low Resource Indic langu…

200 papers

Automatic speech recognition (ASR) for African languages remains constrained by limited labeled data and the lack of systematic guidance on model selection, data scaling, and decoding strategies. Large pre-trained systems such as Whisper,…

There are several domains that own corresponding widely used feature extractors, such as ResNet, BERT, and GPT-x. These models are usually pre-trained on large amounts of unlabeled data by self-supervision and can be effectively applied to…

Computation and Language · Computer Science 2021-01-19 Cheng Yi , Jianzhong Wang , Ning Cheng , Shiyu Zhou , Bo Xu

Sign languages are the primary means of communication for many hard-of-hearing people worldwide. Recently, to bridge the communication gap between the hard-of-hearing community and the rest of the population, several sign language…

Computation and Language · Computer Science 2023-07-12 Abhinav Joshi , Susmit Agrawal , Ashutosh Modi

India is a diverse society with unique challenges in developing AI systems, including linguistic diversity, oral traditions, data accessibility, and scalability. Existing foundation models are primarily trained on English, limiting their…

More than 7,000 known languages are spoken around the world. However, due to the lack of annotated resources, only a small fraction of them are currently covered by speech technologies. Albeit self-supervised speech representations, recent…

Computer Vision and Pattern Recognition · Computer Science 2024-02-21 José-M. Acosta-Triana , David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

We present a culturally-grounded multimodal dataset of 1,060 traditional recipes crowdsourced from rural communities across remote regions of Eastern India, spanning 10 endangered languages. These recipes, rich in linguistic and cultural…

Developing benchmark datasets for low-resource languages poses significant challenges, primarily due to the limited availability of native linguistic experts and the substantial time and cost involved in annotation. Given these challenges,…

Computation and Language · Computer Science 2025-10-28 Rahul Ranjan , Mahendra Kumar Gurve , Anuj , Nitin , Yamuna Prasad

While Indic NLP has made rapid advances recently in terms of the availability of corpora and pre-trained models, benchmark datasets on standard NLU tasks are limited. To this end, we introduce IndicXNLI, an NLI dataset for 11 Indic…

Computation and Language · Computer Science 2022-04-20 Divyanshu Aggarwal , Vivek Gupta , Anoop Kunchukuttan

India's linguistic diversity presents both opportunities and challenges for fintech platforms. While the country has 31 major languages and over 100 minor ones, only 10\% of the population understands English, creating barriers to financial…

Computation and Language · Computer Science 2025-12-02 Bharatdeep Hazarika , Arya Suneesh , Prasanna Devadiga , Pawan Kumar Rajpoot , Anshuman B Suresh , Ahmed Ifthaquar Hussain

Multimodal research has predominantly focused on single-image reasoning, with limited exploration of multi-image scenarios. Recent models have sought to enhance multi-image understanding through large-scale pretraining on interleaved…

Computation and Language · Computer Science 2026-03-26 Shaharukh Khan , Ali Faraz , Abhinav Ravi , Mohd Nauman , Mohd Sarfraz , Akshat Patidar , Raja Kolla , Chandra Khatri , Shubham Agarwal

We conduct an empirical study of cross-lingual transfer using spontaneous, noisy, and code-mixed speech across a wide range of Indic dialects and language varieties. Our results indicate that although ASR performance is generally improved…

Computation and Language · Computer Science 2026-02-12 Akriti Dhasmana , Aarohi Srivastava , David Chiang

We release Rasa, the first multilingual expressive TTS dataset for any Indian language, which contains 10 hours of neutral speech and 1-3 hours of expressive speech for each of the 6 Ekman emotions covering 3 languages: Assamese, Bengali, &…

Computation and Language · Computer Science 2024-09-04 Praveen Srinivasa Varadhan , Ashwin Sankar , Giri Raju , Mitesh M. Khapra

Automatic speech recognition systems have undoubtedly advanced with the integration of multilingual and multitask models such as Whisper, which have shown a promising ability to understand and process speech across a wide range of…

Computation and Language · Computer Science 2025-04-14 Xabier de Zuazo , Eva Navas , Ibon Saratxaga , Inma Hernáez Rioja

With nearly 1.5 billion people and more than 120 major languages, India represents one of the most diverse regions in the world. As multilingual Vision-Language Models (VLMs) gain prominence, robust evaluation methodologies are essential to…

Small Language Models (SLMs) offer efficient alternatives to LLMs for specific domains. The 2023 TinyStories study developed an English dataset that allows SLMs with 1 to 10 million parameters to produce coherent outputs. Our research…

The problem of audio-to-text alignment has seen significant amount of research using complete supervision during training. However, this is typically not in the context of long audio recordings wherein the text being queried does not appear…

Computation and Language · Computer Science 2023-10-11 Piyush Singh Pasi , Karthikeya Battepati , Preethi Jyothi , Ganesh Ramakrishnan , Tanmay Mahapatra , Manoj Singh

Speech recognition is a fascinating process that offers the opportunity to interact and command the machine in the field of human-computer interactions. Speech recognition is a language-dependent system constructed directly based on the…

Computation and Language · Computer Science 2021-09-28 M. F. Mridha , Abu Quwsar Ohi , Md. Abdul Hamid , Muhammad Mostafa Monowar

Automatic Speech Recognition (ASR) generates text which is most of the times devoid of any punctuation. Absence of punctuation is text can affect readability. Also, down stream NLP tasks such as sentiment analysis, machine translation,…

Computation and Language · Computer Science 2022-04-01 Anirudh Gupta , Neeraj Chhimwal , Ankur Dhuriya , Rishabh Gaur , Priyanshi Shah , Harveen Singh Chadha , Vivek Raghavan

Approaching Speech-to-Text and Automatic Speech Recognition problems in low-resource languages is notoriously challenging due to the scarcity of validated datasets and the diversity of dialects. Arabic, Russian, and Portuguese exemplify…

Computation and Language · Computer Science 2025-01-03 Or Haim Anidjar , Revital Marbel , Roi Yozevitch

Developing automatic speech recognition (ASR) systems for low-resource languages is hindered by the scarcity of transcribed corpora. This proof-of-concept study explores songs as an unconventional yet promising data source for Kazakh ASR.…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-10 Rustem Yeshpanov