English
Related papers

Related papers: Crowdsourcing Lexical Diversity

200 papers

Crowdsourcing has been widely used to efficiently obtain labeled datasets for supervised learning from large numbers of human resources at low cost. However, one of the technical challenges in obtaining high-quality results from…

Human-Computer Interaction · Computer Science 2023-02-28 Ryosuke Ueda , Koh Takeuchi , Hisashi Kashima

Lexical resources are crucial for cross-linguistic analysis and can provide new insights into computational models for natural language learning. Here, we present an advanced database for comparative studies of words with multiple meanings,…

Computation and Language · Computer Science 2025-08-22 Annika Tjuka , Robert Forkel , Christoph Rzymski , Johann-Mattis List

Multi-language recipe personalisation and recommendation is an under-explored field of information retrieval in academic and production systems. The existing gaps in our current understanding are numerous, even on fundamental questions such…

Information Retrieval · Computer Science 2020-08-19 Niall Twomey , Mikhail Fain , Andrey Ponikar , Nadine Sarraf

Multilingual large language models (LLMs) have gained prominence, but concerns arise regarding their reliability beyond English. This study addresses the gap in cross-lingual semantic evaluation by introducing a novel benchmark for…

Computation and Language · Computer Science 2024-10-31 Samuel Cahyawijaya , Ruochen Zhang , Holy Lovenia , Jan Christian Blaise Cruz , Elisa Gilbert , Hiroki Nomoto , Alham Fikri Aji

A multilingual collection may contain useful knowledge in other languages to supplement and correct the facts in the original language for Retrieval-Augmented Generation (RAG). However, the vanilla approach that simply concatenates multiple…

Computation and Language · Computer Science 2026-04-30 Rui Qi , Fengran Mo , Sijin Lu , Yufeng Chen , Jian-Yun Nie , Kaiyu Huang

Retrieval-augmented generation (RAG) has become a cornerstone of contemporary NLP, enhancing large language models (LLMs) by allowing them to access richer factual contexts through in-context retrieval. While effective in monolingual…

Computation and Language · Computer Science 2026-03-31 Leonardo Ranaldi , Barry Haddow , Alexandra Birch

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on standard computer vision benchmarks, many of which, however,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Thao Nguyen , Matthew Wallingford , Sebastin Santy , Wei-Chiu Ma , Sewoong Oh , Ludwig Schmidt , Pang Wei Koh , Ranjay Krishna

Typologically diverse benchmarks are increasingly created to track the progress achieved in multilingual NLP. Linguistic diversity of these data sets is typically measured as the number of languages or language families included in the…

Computation and Language · Computer Science 2024-04-17 Tanja Samardzic , Ximena Gutierrez , Christian Bentz , Steven Moran , Olga Pelloni

While generative multilingual models are rapidly being deployed, their safety and fairness evaluations are largely limited to resources collected in English. This is especially problematic for evaluations targeting inherently socio-cultural…

Computation and Language · Computer Science 2024-03-12 Mukul Bhutani , Kevin Robinson , Vinodkumar Prabhakaran , Shachi Dave , Sunipa Dev

Ensuring that the benefits of sign language technologies are distributed equitably among all community members is crucial. Thus, it is important to address potential biases and inequities that may arise from the design or use of these…

Computation and Language · Computer Science 2024-10-08 Katherine Atwell , Danielle Bragg , Malihe Alikhani

Sentiment analysis aims to uncover emotions conveyed through information. In its simplest form, it is performed on a polarity basis, where the goal is to classify information with positive or negative emotion. Recent research has explored…

Computation and Language · Computer Science 2017-10-12 Giannis Haralabopoulos , Elena Simperl

Linguistically diverse datasets are critical for training and evaluating robust machine learning systems, but data collection is a costly process that often requires experts. Crowdsourcing the process of paraphrase generation is an…

Computation and Language · Computer Science 2020-06-05 Youxuan Jiang , Jonathan K. Kummerfeld , Walter S. Lasecki

How good are humans at writing and judging responses in retrieval-augmented generation (RAG) scenarios? To answer this question, we investigate the efficacy of crowdsourcing for RAG through two complementary studies: response writing and…

Information Retrieval · Computer Science 2025-04-23 Lukas Gienapp , Tim Hagen , Maik Fröbe , Matthias Hagen , Benno Stein , Martin Potthast , Harrisen Scells

Commonsense reasoning (CSR) requires the model to be equipped with general world knowledge. While CSR is a language-agnostic process, most comprehensive knowledge sources are in few popular languages, especially English. Thus, it remains…

Computation and Language · Computer Science 2021-10-19 Yuwei Fang , Shuohang Wang , Yichong Xu , Ruochen Xu , Siqi Sun , Chenguang Zhu , Michael Zeng

Large language models (LLMs) tend to generate homogenous texts, which may impact the diversity of knowledge generated across different outputs. Given their potential to replace existing forms of knowledge acquisition, this poses a risk of…

Data scarcity in low-resource languages can be addressed with word-to-word translations from labeled task data in high-resource languages using bilingual lexicons. However, bilingual lexicons often have limited lexical overlap with task…

Computation and Language · Computer Science 2024-10-29 Zheng-Xin Yong , Cristina Menghini , Stephen H. Bach

Distant supervision is a popular method for performing relation extraction from text that is known to produce noisy labels. Most progress in relation extraction and classification has been made with crowdsourced corrections to…

Computation and Language · Computer Science 2022-09-21 Anca Dumitrache , Lora Aroyo , Chris Welty

We introduce Multi-SimLex, a large-scale lexical resource and evaluation benchmark covering datasets for 12 typologically diverse languages, including major languages (e.g., Mandarin Chinese, Spanish, Russian) as well as less-resourced ones…

Cross-lingual information retrieval (CLIR) addresses the challenge of retrieving relevant documents written in languages different from that of the original query. Research in this area has typically framed the task as monolingual retrieval…

Information Retrieval · Computer Science 2025-10-02 Roksana Goworek , Olivia Macmillan-Scott , Eda B. Özyiğit

This paper explores grading text-based audio retrieval relevances with crowdsourcing assessments. Given a free-form text (e.g., a caption) as a query, crowdworkers are asked to grade audio clips using numeric scores (between 0 and 100) to…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-16 Huang Xie , Khazar Khorrami , Okko Räsänen , Tuomas Virtanen