English
Related papers

Related papers: Building and curating conversational corpora for d…

200 papers

India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognized as scheduled languages in the Indian Constitution. Despite…

We present the Claire French Dialogue Dataset (CFDD), a resource created by members of LINAGORA Labs in the context of the OpenLLM France initiative. CFDD is a corpus containing roughly 160 million words from transcripts and stage plays in…

Computation and Language · Computer Science 2023-11-29 Julie Hunter , Jérôme Louradour , Virgile Rennard , Ismaïl Harrando , Guokan Shang , Jean-Pierre Lorré

Research in speech technologies and comparative linguistics depends on access to diverse and accessible speech data. The UCLA Phonetics Lab Archive is one of the earliest multilingual speech corpora, with long-form audio recordings and…

Computation and Language · Computer Science 2024-03-29 Eleanor Chodroff , Blaž Pažon , Annie Baker , Steven Moran

Large Language Models (LLMs) have demonstrated remarkable performance across various disciplines and tasks. However, benchmarking their capabilities with multilingual spoken queries remains largely unexplored. In this study, we introduce…

Computation and Language · Computer Science 2025-05-27 Firoj Alam , Md Arid Hasan , Shammur Absar Chowdhury

This survey focuses in encoder Language Models for solving tasks in the clinical domain in the Spanish language. We review the contributions of 17 corpora focused mainly in clinical tasks, then list the most relevant Spanish Language Models…

Computation and Language · Computer Science 2023-08-07 Guillem García Subies , Álvaro Barbero Jiménez , Paloma Martínez Fernández

We present our first efforts towards building a single multilingual automatic speech recognition (ASR) system that can process code-switching (CS) speech in five languages spoken within the same population. This contrasts with related prior…

Computation and Language · Computer Science 2018-07-31 Emre Yılmaz , Astik Biswas , Ewald van der Westhuizen , Febe de Wet , Thomas Niesler

MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual dataset we have built for the WSDM 2023 Cup challenge that focuses on ad hoc retrieval across 18 different languages, which collectively encompass…

The effectiveness of Large Language Models (LLMs) depends heavily on the availability of high-quality post-training data, particularly instruction-tuning and preference-based examples. Existing open-source datasets, however, often lack…

Computation and Language · Computer Science 2025-10-09 Neel Prabhanjan Rachamalla , Aravind Konakalla , Gautam Rajeev , Ashish Kulkarni , Chandra Khatri , Shubham Agarwal

Datasets are foundational to many breakthroughs in modern artificial intelligence. Many recent achievements in the space of natural language processing (NLP) can be attributed to the finetuning of pre-trained models on a diverse set of…

Parallel corpora play an important role in training machine translation (MT) models, particularly for low-resource languages where high-quality bilingual data is scarce. This review provides a comprehensive overview of available parallel…

Computation and Language · Computer Science 2025-04-23 Rahul Raja , Arpita Vats

We present "Testimole-conversational" a massive collection of discussion boards messages in the Italian language. The large size of the corpus, more than 30B word-tokens (1996-2024), renders it an ideal dataset for native Italian Large…

Computation and Language · Computer Science 2026-04-10 Matteo Rinaldi , Rossella Varvara , Viviana Patti

Various formal languages have been proposed in the literature for the individual-based modelling of ecological systems. These languages differ in their treatment of time and space. Each modelling language offers a distinct view and…

Logic in Computer Science · Computer Science 2019-01-31 Mauricio Toro

Evaluating English ASR systems for conversational AI applications remains difficult, as many publicly available corpora are either pre-segmented into short segments, consist of read or prepared speech, or lack explicit dialect annotations…

Computation and Language · Computer Science 2026-05-01 Eugen Beck , Sarah Beranek , Uma Moothiringote , Daniel Mann , Wilfried Michel , Katie Nguyen , Taylor Tragemann

With the emergence of increasingly powerful large language models, there is a burgeoning interest in leveraging these models for casual conversation and role-play applications. However, existing conversational and role-playing datasets…

Computation and Language · Computer Science 2023-08-14 Tear Gosling , Alpin Dale , Yinhe Zheng

Recently, multilingual artificial intelligence assistants, exemplified by ChatGPT, have gained immense popularity. As a crucial gateway to human-computer interaction, multilingual automatic speech recognition (ASR) has also garnered…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-27 Song Li , Yongbin You , Xuezhi Wang , Zhengkun Tian , Ke Ding , Guanglu Wan

The driving factors behind the development of large language models (LLMs) with impressive learning capabilities are their colossal model sizes and extensive training datasets. Along with the progress in natural language processing, LLMs…

Computation and Language · Computer Science 2023-09-19 Thuat Nguyen , Chien Van Nguyen , Viet Dac Lai , Hieu Man , Nghia Trung Ngo , Franck Dernoncourt , Ryan A. Rossi , Thien Huu Nguyen

Automatic language identification is frequently framed as a multi-class classification problem. However, when creating digital corpora for less commonly written languages, it may be more appropriate to consider it a data mining problem. For…

Computation and Language · Computer Science 2025-03-11 Rasul Dent , Pedro Ortiz Suarez , Thibault Clérice , Benoît Sagot

A conversation corpus is essential to build interactive AI applications. However, the demographic information of the participants in such corpora is largely underexplored mainly due to the lack of individual data in many corpora. In this…

Computation and Language · Computer Science 2022-04-21 Haewoon Kwak , Jisun An , Kunwoo Park

Aligning large language models (LLMs) with human preferences has proven to drastically improve usability and has driven rapid adoption as demonstrated by ChatGPT. Alignment techniques such as supervised fine-tuning (SFT) and reinforcement…

The BabyLM Challenge is a community effort to close the data-efficiency gap between human and computational language learners. Participants compete to optimize language model training on a fixed language data budget of 100 million words or…

‹ Prev 1 3 4 5 6 7 10 Next ›