English
Related papers

Related papers: Documenting Geographically and Contextually Divers…

200 papers

Democratizing access to natural language processing (NLP) technology is crucial, especially for underrepresented and extremely low-resource languages. Previous research has focused on developing labeled and unlabeled corpora for these…

How can large language models (LLMs) serve users with varying preferences that may conflict across cultural, political, or other dimensions? To advance this challenge, this paper establishes four key results. First, we demonstrate, through…

The paper proposes an approach to transcend multicultural and multilingual barriers in the use and reuse of geographical data at the European level. The approach aims at sharing scientific terms in the field of nature conservation with the…

Digital Libraries · Computer Science 2011-07-11 Monica De Martino , Riccardo Albertoni

Large language models are typically trained by treating text as a single global distribution, often resulting in geographically homogenized behavior. We study metadata conditioning as a lightweight approach for localization, pre-training 31…

Computation and Language · Computer Science 2026-01-22 Anjishnu Mukherjee , Ziwei Zhu , Antonios Anastasopoulos

In this project, we semantically enriched and enhanced the metadata of long text documents, theses and dissertations, retrieved from the HathiTrust Digital Library in English published from 1920 to 2020 through a combination of manual…

Digital Libraries · Computer Science 2025-06-27 Manika Lamba , You Peng , Sophie Nikolov , Glen Layne-Worthey , J. Stephen Downie

Translation between natural language and source code can help software development by enabling developers to comprehend, ideate, search, and write computer programs in natural language. Despite growing interest from the industry and the…

Data originating from the Web, sensor readings and social media result in increasingly huge datasets. The so called Big Data comes with new scientific and technological challenges while creating new opportunities, hence the increasing…

Artificial Intelligence · Computer Science 2020-02-19 Ilias Tachmazidis , Grigoris Antoniou , Wolfgang Faber

Maintaining literature databases and online bibliographies is a core responsibility of metadata aggregators such as digital libraries. In the process of monitoring all the available data sources the question arises which data source should…

Digital Libraries · Computer Science 2018-04-18 Mandy Neumann , Christopher Michels , Philipp Schaer , Ralf Schenkel

High-resource languages such as English, enables the pretraining of high-quality large language models (LLMs). The same can not be said for most other languages as LLMs still underperform for non-English languages, likely due to a gap in…

Computation and Language · Computer Science 2025-02-20 Jiayi Wang , Yao Lu , Maurice Weber , Max Ryabinin , David Adelani , Yihong Chen , Raphael Tang , Pontus Stenetorp

Recent breakthroughs in large language modeling have facilitated rigorous exploration of their application in diverse tasks related to tabular data modeling, such as prediction, tabular data synthesis, question answering, and table…

Writing a survey paper on one research topic usually needs to cover the salient content from numerous related papers, which can be modeled as a multi-document summarization (MDS) task. Existing MDS datasets usually focus on producing the…

Computation and Language · Computer Science 2023-02-10 Shuaiqi Liu , Jiannong Cao , Ruosong Yang , Zhiyuan Wen

Multilingual instruction fine-tuning (IFT) empowers large language models to generalize across diverse linguistic and cultural contexts; however, high-quality, systematically curated multilingual IFT datasets remain scarce. To address this…

Computation and Language · Computer Science 2026-05-01 Chunguang Zhao , Yilun Liu , Pufan Zeng , Yuanchang Luo , Shimin Tao , Minggui He , Weibin Meng , Song Xu , Chen Liu , Hongxia Ma , Li Zhang , Boxing Chen , Daimeng Wei

Modern machine learning research relies on relatively few carefully curated datasets. Even in these datasets, and typically in `untidy' or raw data, practitioners are faced with significant issues of data quality and diversity which can be…

Machine Learning · Computer Science 2022-09-22 Shoaib Ahmed Siddiqui , Nitarshan Rajkumar , Tegan Maharaj , David Krueger , Sara Hooker

Combating online hate speech in multilingual settings requires approaches that go beyond English-centric models and capture the cultural and linguistic diversity of global online discourse. This paper presents a comprehensive survey and…

Computation and Language · Computer Science 2026-03-23 Zahra Safdari Fesaghandis , Suman Kalyan Maity

Perceptions of hate can vary greatly across cultural contexts. Hate speech (HS) datasets, however, have traditionally been developed by language. This hides potential cultural biases, as one language may be spoken in different countries…

Computation and Language · Computer Science 2025-05-20 Manuel Tonneau , Diyi Liu , Samuel Fraiberger , Ralph Schroeder , Scott A. Hale , Paul Röttger

This paper presents an overview of a program designed to address the growing need for developing freely available speech resources for under-represented languages. At present we have released 38 datasets for building text-to-speech and…

This paper traces strategies used by the Radio Preservation Task Force of the Library of Congress's National Recording Preservation Board to develop a publicly searchable database documenting extant radio materials held by collecting…

Digital Libraries · Computer Science 2020-03-04 Emily Goodmann , Mark A. Matienzo , Shawn VanCour , William Vanden Dries

Large language models' (LLMs) abilities are drawn from their pretraining data, and model development begins with data curation. However, decisions around what data is retained or removed during this initial stage are under-scrutinized. In…

Computation and Language · Computer Science 2024-06-24 Li Lucy , Suchin Gururangan , Luca Soldaini , Emma Strubell , David Bamman , Lauren F. Klein , Jesse Dodge

This project addresses the challenges of responsible and fair resource allocation in data science (DS), focusing on DS queries evaluation. Current DS practices often overlook the broader socio-economic, environmental, and ethical…

Databases · Computer Science 2025-02-18 Genoveva Vargas-Solar