English
Related papers

Related papers: Google Crowdsourced Speech Corpora and Related Ope…

200 papers

Conversational agents are gaining popularity with the increasing ubiquity of smart devices. However, training agents in a data driven manner is challenging due to a lack of suitable corpora. This paper presents a novel method for gathering…

Computation and Language · Computer Science 2018-09-20 Joachim Fainberg , Ben Krause , Mihai Dobre , Marco Damonte , Emmanuel Kahembwe , Daniel Duma , Bonnie Webber , Federico Fancellu

Korean is often referred to as a low-resource language in the research community. While this claim is partially true, it is also because the availability of resources is inadequately advertised and curated. This work curates and reviews a…

Computation and Language · Computer Science 2023-05-17 Won Ik Cho , Sangwhan Moon , Youngsook Song

This paper briefly reports our ongoing attempt at the development of a multi-platform browser-based speech recording system. We designed the system toward a service of providing open service of building large-scale speech corpora at a…

Human-Computer Interaction · Computer Science 2019-12-20 Keita Ishizuka , Takashi Nose

This tutorial (https://tum-nlp.github.io/low-resource-tutorial) is designed for NLP practitioners, researchers, and developers working with multilingual and low-resource languages who seek to create more equitable and socially impactful…

Computation and Language · Computer Science 2025-12-17 Ekaterina Artemova , Laurie Burchell , Daryna Dementieva , Shu Okabe , Mariya Shmatova , Pedro Ortiz Suarez

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other…

This preprint describes work in progress on LR-Sum, a new permissively-licensed dataset created with the goal of enabling further research in automatic summarization for less-resourced languages. LR-Sum contains human-written summaries for…

Computation and Language · Computer Science 2023-10-30 Chester Palen-Michel , Constantine Lignos

Democratizing access to natural language processing (NLP) technology is crucial, especially for underrepresented and extremely low-resource languages. Previous research has focused on developing labeled and unlabeled corpora for these…

Most speech and language technologies are trained with massive amounts of speech and text information. However, most of the world languages do not have such resources or stable orthography. Systems constructed under these almost zero…

This study explores methods to increase data volume for low-resource languages using techniques such as crowdsourcing, pseudo-labeling, advanced data preprocessing and various permissive data sources such as audiobooks, Common Voice,…

We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx talks in 8 source…

Computation and Language · Computer Science 2021-06-16 Elizabeth Salesky , Matthew Wiesner , Jacob Bremerman , Roldano Cattoni , Matteo Negri , Marco Turchi , Douglas W. Oard , Matt Post

With the advent of generative audio features, there is an increasing need for rapid evaluation of their impact on speech intelligibility. Beyond the existing laboratory measures, which are expensive and do not scale well, there has been…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-25 Laura Lechler , Kamil Wojcicki

Advances in speech and language technologies enable tools such as voice-search, text-to-speech, speech recognition and machine translation. These are however only available for high resource languages like English, French or Chinese.…

Modern machine learning models for audio tasks often exhibit superior performance on English and other well-resourced languages, primarily due to the abundance of available training data. This disparity leads to an unfair performance gap…

Computation and Language · Computer Science 2025-11-26 Wesley Bian , Xiaofeng Lin , Guang Cheng

We present a survey covering the state of the art in low-resource machine translation research. There are currently around 7000 languages spoken in the world and almost all language pairs lack significant resources for training machine…

Computation and Language · Computer Science 2022-02-08 Barry Haddow , Rachel Bawden , Antonio Valerio Miceli Barone , Jindřich Helcl , Alexandra Birch

Modern speech synthesis techniques can produce natural-sounding speech given sufficient high-quality data and compute resources. However, such data is not readily available for many languages. This paper focuses on speech synthesis for…

Computation and Language · Computer Science 2022-07-05 Perez Ogayo , Graham Neubig , Alan W Black

With the rise of artificial intelligence (AI) and the growing use of deep-learning architectures, the question of ethics, transparency and fairness of AI systems has become a central concern within the research community. We address…

Computation and Language · Computer Science 2020-03-19 Mahault Garnerin , Solange Rossato , Laurent Besacier

This paper presents the creation of initial bilingual corpora for thirteen very low-resource languages of India, all from Northeast India. It also presents the results of initial translation efforts in these languages. It creates the…

Computation and Language · Computer Science 2023-12-11 Atnafu Lambebo Tonja , Melkamu Mersha , Ananya Kalita , Olga Kolesnikova , Jugal Kalita

Natural Language Processing (NLP) for low-resource languages remains fundamentally constrained by the lack of textual corpora, standardized orthographies, and scalable annotation pipelines. While recent advances in large language models…

Computation and Language · Computer Science 2026-02-10 Bonaventure F. P. Dossou , Henri Aïdasso

Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets. Nevertheless, current datasets cover a limited number of languages. With the aim to foster research in…

Computation and Language · Computer Science 2020-10-27 Changhan Wang , Anne Wu , Juan Pino
‹ Prev 1 2 3 10 Next ›