English
Related papers

Related papers: Mapping Languages: The Corpus of Global Language U…

200 papers

The lack of large-scale datasets has been a major hindrance to the development of NLP tasks such as spelling correction and grammatical error correction (GEC). As a complementary new resource for these tasks, we present the GitHub Typo…

Computation and Language · Computer Science 2019-12-02 Masato Hagiwara , Masato Mita

Word choice is dependent on the cultural context of writers and their subjects. Different words are used to describe similar actions, objects, and features based on factors such as class, race, gender, geography and political affinity.…

Computation and Language · Computer Science 2018-07-02 Taylor Arnold , Lauren Tilton

We introduce RadioTalk, a corpus of speech recognition transcripts sampled from talk radio broadcasts in the United States between October of 2018 and March of 2019. The corpus is intended for use by researchers in the fields of natural…

Computation and Language · Computer Science 2019-09-18 Doug Beeferman , William Brannon , Deb Roy

We present a novel corpus of 445 human- and computer-generated documents, comprising about 27,000 clauses, annotated for semantic clause types and coherence relations that allow for nuanced comparison of artificial and natural discourse…

We present an analysis pipeline and best practice guidelines for building and curating corpora of everyday conversation in diverse languages. Surveying language documentation corpora and other resources that cover 67 languages and varieties…

Computation and Language · Computer Science 2022-05-11 Andreas Liesenfeld , Mark Dingemanse

This paper introduces the Ubuntu Dialogue Corpus, a dataset containing almost 1 million multi-turn dialogues, with a total of over 7 million utterances and 100 million words. This provides a unique resource for research into building…

Computation and Language · Computer Science 2016-07-26 Ryan Lowe , Nissan Pow , Iulian Serban , Joelle Pineau

ArabJobs is a publicly available corpus of Arabic job advertisements collected from Egypt, Jordan, Saudi Arabia, and the United Arab Emirates. Comprising over 8,500 postings and more than 550,000 words, the dataset captures linguistic,…

Computation and Language · Computer Science 2025-09-29 Mo El-Haj

Translations capture important information about languages that can be used as implicit supervision in learning linguistic properties and semantic representations. In an information-centric view, translated texts may be considered as…

Computation and Language · Computer Science 2018-02-02 Jörg Tiedemann

Cultural diversity encoded within languages of the world is at risk, as many languages have become endangered in the last decades in a context of growing globalization. To preserve this diversity, it is first necessary to understand what…

Physics and Society · Physics 2022-10-10 Thomas Louf , David Sanchez , Jose J. Ramasco

This paper introduces Multilingual LibriSpeech (MLS) dataset, a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages, including about 44.5K hours of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-22 Vineel Pratap , Qiantong Xu , Anuroop Sriram , Gabriel Synnaeve , Ronan Collobert

Natural Language Processing (NLP) is increasingly used as a key ingredient in critical decision-making systems such as resume parsers used in sorting a list of job candidates. NLP systems often ingest large corpora of human text, attempting…

Computation and Language · Computer Science 2020-07-14 Esma Wali , Yan Chen , Christopher Mahoney , Thomas Middleton , Marzieh Babaeianjelodar , Mariama Njie , Jeanna Neefe Matthews

Cooperative video games, where multiple participants must coordinate by communicating and reasoning under uncertainty in complex environments, yield a rich source of language data. We collect the Portal Dialogue Corpus: a corpus of 11.5…

Arabic is a linguistically and culturally rich language with a vast vocabulary that spans scientific, religious, and literary domains. Yet, large-scale lexical datasets linking Arabic words to precise definitions remain limited. We present…

Computation and Language · Computer Science 2026-01-30 Serry Sibaee , Yasser Alhabashi , Nadia Sibai , Yara Farouk , Adel Ammar , Sawsan AlHalawani , Wadii Boulila

Existing research on fairness evaluation of document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes. In this work, we assemble and publish a multilingual Twitter corpus…

Computation and Language · Computer Science 2020-03-04 Xiaolei Huang , Linzi Xing , Franck Dernoncourt , Michael J. Paul

In recent years, large language models (e.g., Open AI's GPT-4, Meta's LLaMa, Google's PaLM) have become the dominant approach for building AI systems to analyze and generate language online. However, the automated systems that increasingly…

Computation and Language · Computer Science 2023-06-14 Gabriel Nicholas , Aliya Bhatia

This paper evaluates global-scale dialect identification for 14 national varieties of English as a means for studying syntactic variation. The paper makes three main contributions: (i) introducing data-driven language mapping as a method…

Computation and Language · Computer Science 2019-04-12 Jonathan Dunn

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on standard computer vision benchmarks, many of which, however,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Thao Nguyen , Matthew Wallingford , Sebastin Santy , Wei-Chiu Ma , Sewoong Oh , Ludwig Schmidt , Pang Wei Koh , Ranjay Krishna

Providing better language tools for low-resource and endangered languages is imperative for equitable growth. Recent progress with massively multilingual pretrained models has proven surprisingly effective at performing zero-shot transfer…

Computation and Language · Computer Science 2022-11-10 Louis Clouâtre , Prasanna Parthasarathi , Amal Zouaq , Sarath Chandar

The need for a comprehensive study to explore various aspects of online social media has been instigated by many researchers. This paper gives an insight into the social platform, Twitter. In this present work, we have illustrated stepwise…

Social and Information Networks · Computer Science 2021-05-24 Shahab Saquib Sohail , Mohammad Muzammil Khan , M. Afshar Alam

This paper presents StoryDB - a broad multi-language dataset of narratives. StoryDB is a corpus of texts that includes stories in 42 different languages. Every language includes 500+ stories. Some of the languages include more than 20 000…

Computation and Language · Computer Science 2022-11-15 Alexey Tikhonov , Igor Samenko , Ivan P. Yamshchikov