English
Related papers

Related papers: Documenting Geographically and Contextually Divers…

200 papers

Large Language Models (LLMs) have demonstrated immense potential in artificial intelligence across various domains, including healthcare. However, their efficacy is hindered by the need for high-quality labeled data, which is often…

Computation and Language · Computer Science 2024-05-24 P. Barai , G. Leroy , P. Bisht , J. M. Rothman , S. Lee , J. Andrews , S. A. Rice , A. Ahmed

Deep learning-based approaches for automatic document layout analysis and content extraction have the potential to unlock rich information trapped in historical documents on a large scale. One major hurdle is the lack of large datasets for…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Zejiang Shen , Kaixuan Zhang , Melissa Dell

Information retrieval (IR) is the task of finding relevant documents in response to a user query. Although Spanish is the second most spoken native language, there are few Spanish IR datasets, which limits the development of information…

Computation and Language · Computer Science 2025-11-20 Francisco Valentini , Viviana Cotik , Damián Furman , Ivan Bercovich , Edgar Altszyler , Juan Manuel Pérez

Data scarcity has been a long standing issue in the field of open-domain social dialogue. To quench this thirst, we present SODA: the first publicly available, million-scale high-quality social dialogue dataset. By contextualizing social…

Computation and Language · Computer Science 2023-10-25 Hyunwoo Kim , Jack Hessel , Liwei Jiang , Peter West , Ximing Lu , Youngjae Yu , Pei Zhou , Ronan Le Bras , Malihe Alikhani , Gunhee Kim , Maarten Sap , Yejin Choi

Low-quality data can cause downstream problems in high-stakes applications. Data-centric approach emphasizes on improving dataset quality to enhance model performance. High-quality datasets are needed for general-purpose Large Language…

Computation and Language · Computer Science 2023-10-13 Iva Bojic , Josef Halim , Verena Suharman , Sreeja Tar , Qi Chwen Ong , Duy Phung , Mathieu Ravaut , Shafiq Joty , Josip Car

Machine Translation is a mature technology for many high-resource language pairs. However in the context of low-resource languages, there is a paucity of parallel data datasets available for developing translation models. Furthermore, the…

Computation and Language · Computer Science 2024-03-07 Séamus Lankford , Haithem Afli , Órla Ní Loinsigh , Andy Way

Large language models have shown unprecedented abilities in generating linguistically coherent and syntactically correct natural language output. However, they often return incorrect and inconsistent answers to input questions. Due to the…

Databases · Computer Science 2023-12-27 Jasmin Mousavi , Arash Termehchy

Current research on hate speech analysis is typically oriented towards monolingual and single classification tasks. In this paper, we present a new multilingual hate speech analysis dataset for English, Hindi, Arabic, French, German and…

Computation and Language · Computer Science 2023-04-04 Ankit Yadav , Shubham Chandel , Sushant Chatufale , Anil Bandhakavi

Large language models are trained on massive scrapes of the web, as required by current scaling laws. Most progress is made for English, given its abundance of high-quality pretraining data. For most other languages, however, such high…

Computation and Language · Computer Science 2025-02-07 Skyler Seto , Maartje ter Hoeve , Richard He Bai , Natalie Schluter , David Grangier

Although numerous datasets have been developed to support dialogue systems, most existing chit-chat datasets overlook the cultural nuances inherent in natural human conversations. To address this gap, we introduce SEADialogues, a culturally…

Advances in data analytics bring with them civil rights implications. Data-driven and algorithmic decision making increasingly determine how businesses target advertisements to consumers, how police departments monitor individuals or…

Computers and Society · Computer Science 2017-06-13 Solon Barocas , Elizabeth Bradley , Vasant Honavar , Foster Provost

High-quality data resources play a crucial role in learning large language models (LLMs), particularly for low-resource languages like Cantonese. Despite having more than 85 million native speakers, Cantonese is still considered a…

Computation and Language · Computer Science 2025-03-06 Jiyue Jiang , Alfred Kar Yin Truong , Yanyu Chen , Qinghang Bao , Sheng Wang , Pengan Chen , Jiuming Wang , Lingpeng Kong , Yu Li , Chuan Wu

Large language models have recently advanced the state of the art on many natural language processing benchmarks. The newest generation of models can be applied to a variety of tasks with little to no specialized training. This technology…

Databases · Computer Science 2023-06-16 Immanuel Trummer

Traditionally, large language models have been either trained on general web crawls or domain-specific data. However, recent successes of generative large language models, have shed light on the benefits of cross-domain datasets. To examine…

Multilingual language models have been a crucial breakthrough as they considerably reduce the need of data for under-resourced languages. Nevertheless, the superiority of language-specific models has already been proven for languages having…

Due to the black-box nature of large language models (LLMs) and the realism of their generated content, issues such as hallucinations, bias, unfairness, and copyright infringement have become significant. In this context, sourcing…

Computation and Language · Computer Science 2026-01-01 Liang Pang , Jia Gu , Sunhao Dai , Zihao Wei , Zenghao Duan , Kangxi Wu , Zhiyi Yin , Jun Xu , Huawei Shen , Xueqi Cheng

Large language models (LLMs) with extended context windows enable tasks requiring extensive information integration but are limited by the scarcity of high-quality, diverse datasets for long-context instruction tuning. Existing data…

Computation and Language · Computer Science 2025-02-25 Jiaxi Li , Xingxing Zhang , Xun Wang , Xiaolong Huang , Li Dong , Liang Wang , Si-Qing Chen , Wei Lu , Furu Wei

Availability, collection and access to quantitative data, as well as its limitations, often make qualitative data the resource upon which development programs heavily rely. Both traditional interview data and social media analysis can…

Computation and Language · Computer Science 2017-09-19 Philipp Broniecki , Anna Hanchar , Slava J. Mikhaylov

Query-based document summarization aims to extract or generate a summary of a document which directly answers or is relevant to the search query. It is an important technique that can be beneficial to a variety of applications such as…

Artificial Intelligence · Computer Science 2020-10-29 Mingjun Zhao , Shengli Yan , Bang Liu , Xinwang Zhong , Qian Hao , Haolan Chen , Di Niu , Bowei Long , Weidong Guo

BOUQuET is a multi-way, multicentric and multi-register/domain dataset and benchmark, and a broader collaborative initiative. This dataset is handcrafted in 8 non-English languages. Each of these source languages are representative of the…