English
Related papers

Related papers: Learnings from Technological Interventions in a Lo…

200 papers

Low-resource languages serve as invaluable repositories of human history, embodying cultural evolution and intellectual diversity. Despite their significance, these languages face critical challenges, including data scarcity and…

Exploiting cognates for transfer learning in under-resourced languages is an exciting opportunity for language understanding tasks, including unsupervised machine translation, named entity recognition and information retrieval. Previous…

Computation and Language · Computer Science 2023-11-10 Koustava Goswami , Priya Rani , Theodorus Fransen , John P. McCrae

The world is passing through a major revolution called the information revolution, in which information and knowledge is becoming available to people in unprecedented amounts wherever and whenever they need it. Those societies which fail to…

Computation and Language · Computer Science 2007-05-23 Akshar Bharati , Rajeev Sangal

In this paper, we offer an overview of indigenous languages, identifying the causes of their devaluation and the need for legislation on language rights. We review the technologies used to revitalize these languages, finding that when they…

Computers and Society · Computer Science 2025-04-03 Silvia Fernandez-Sabido , Laura Peniche-Sabido

The hearing-impaired community in India deserves the access to tools that help them communicate, however, there is limited known technology solutions that make use of Indian Sign Language (ISL) at present. Even though there are many ISL…

Machine Learning · Computer Science 2024-12-11 Smruti Jagtap , Kanika Jadhav , Rushikesh Temkar , Minal Deshmukh

This paper proposes two innovative methodologies to construct customized Common Voice datasets for low-resource languages like Hindi. The first methodology leverages Bark, a transformer-based text-to-audio model developed by Suno, and…

Sound · Computer Science 2024-01-11 Anand Kamble , Aniket Tathe , Suyash Kumbharkar , Atharva Bhandare , Anirban C. Mitra

Spoken Language Identification (LID) is an important sub-task of Automatic Speech Recognition(ASR) that is used to classify the language(s) in an audio segment. Automatic LID plays an useful role in multilingual countries. In various…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-02 Parth Shastri , Chirag Patil , Poorval Wanere , Shrinivas Mahajan , Abhishek Bhatt , Hardik Sailor

Natural Language Processing (NLP) for low-resource languages remains fundamentally constrained by the lack of textual corpora, standardized orthographies, and scalable annotation pipelines. While recent advances in large language models…

Computation and Language · Computer Science 2026-02-10 Bonaventure F. P. Dossou , Henri Aïdasso

The objective of this work is to explore the learning of visually grounded speech models (VGS) from multilingual perspective. Bilingual VGS models are generally trained with an equal number of spoken captions from both languages. However,…

Computation and Language · Computer Science 2023-03-31 Hyeonggon Ryu , Arda Senocak , In So Kweon , Joon Son Chung

In this work, we focus on low-resource dependency parsing for multiple languages. Several strategies are tailored to enhance performance in low-resource scenarios. While these are well-known to the community, it is not trivial to select the…

Computation and Language · Computer Science 2023-01-31 Jivnesh Sandhan , Laxmidhar Behera , Pawan Goyal

Providing better language tools for low-resource and endangered languages is imperative for equitable growth. Recent progress with massively multilingual pretrained models has proven surprisingly effective at performing zero-shot transfer…

Computation and Language · Computer Science 2022-11-10 Louis Clouâtre , Prasanna Parthasarathi , Amal Zouaq , Sarath Chandar

S\'ami, an indigenous language group comprising multiple languages, faces digital marginalization due to the limited availability of data and sophisticated language models designed for its linguistic intricacies. This work focuses on…

Computation and Language · Computer Science 2024-05-10 Ronny Paul , Himanshu Buckchash , Shantipriya Parida , Dilip K. Prasad

Developing culturally grounded multilingual AI systems remains challenging, particularly for low-resource languages. While synthetic data offers promise, its effectiveness in multilingual and multicultural contexts is underexplored. We…

Computation and Language · Computer Science 2026-02-27 Pranjal A. Chitale , Varun Gumma , Sanchit Ahuja , Prashant Kodali , Manan Uppadhyay , Deepthi Sudharsan , Sunayana Sitaram

Low-resource languages often face challenges in acquiring high-quality language data due to the reliance on translation-based methods, which can introduce the translationese effect. This phenomenon results in translated sentences that lack…

Artificial intelligence (AI) is diffusing globally at unprecedented speed, but adoption remains uneven. Frontier Large Language Models (LLMs) are known to perform poorly on low-resource languages due to data scarcity. We hypothesize that…

Computation and Language · Computer Science 2025-11-05 Amit Misra , Syed Waqas Zamir , Wassim Hamidouche , Inbal Becker-Reshef , Juan Lavista Ferres

Online abusive content detection, particularly in low-resource settings and within the audio modality, remains underexplored. We investigate the potential of pre-trained audio representations for detecting abusive language in low-resource…

Computation and Language · Computer Science 2024-12-16 Aditya Narayan Sankaran , Reza Farahbakhsh , Noel Crespi

Existing research in measuring and mitigating gender bias predominantly centers on English, overlooking the intricate challenges posed by non-English languages and the Global South. This paper presents the first comprehensive study delving…

End-to-end text-to-speech (TTS) systems have been developed for European languages like English and Spanish with state-of-the-art speech quality, prosody, and naturalness. However, development of end-to-end TTS for Indian languages is…

Computation and Language · Computer Science 2022-12-08 Ankur Debnath , Shridevi S Patil , Gangotri Nadiger , Ramakrishnan Angarai Ganesan

In this paper, we explore the utility of translationese as synthetic data created using machine translation for pre-training language models (LMs) for low-resource languages (LRLs). Our simple methodology consists of translating large…

Computation and Language · Computer Science 2025-07-08 Meet Doshi , Raj Dabre , Pushpak Bhattacharyya