中文
相关论文

相关论文: Low-resourced Languages and Online Knowledge Repos…

200 篇论文

Training data for machine learning models can come from many different sources, which can be of dubious quality. For resource-rich languages like English, there is a lot of data available, so we can afford to throw out the dubious data. For…

计算与语言 · 计算机科学 2021-03-31 Andrew Zupon , Evan Crew , Sandy Ritchie

This tutorial (https://tum-nlp.github.io/low-resource-tutorial) is designed for NLP practitioners, researchers, and developers working with multilingual and low-resource languages who seek to create more equitable and socially impactful…

计算与语言 · 计算机科学 2025-12-17 Ekaterina Artemova , Laurie Burchell , Daryna Dementieva , Shu Okabe , Mariya Shmatova , Pedro Ortiz Suarez

Language is a form of symbolic capital that affects people's lives in many ways (Bourdieu1977,1991). As a powerful means of communication, it reflects identities, cultures, traditions, and societies more broadly. Therefore, data in a given…

计算与语言 · 计算机科学 2025-06-02 Nedjma Ousidhoum , Meriem Beloucif , Saif M. Mohammad

Data voids--areas of the internet where reliable information is scarce or absent--pose significant challenges to online health information seeking, particularly for users operating in low-web data languages. These voids are increasingly…

We present an open-source online dictionary editing system, Ve'rdd, that offers a chance to re-evaluate and edit grassroots dictionaries that have been exposed to multiple amateur editors. The idea is to incorporate community activities…

计算与语言 · 计算机科学 2020-12-07 Khalid Alnajjar , Mika Hämäläinen , Jack Rueter , Niko Partanen

On Wikipedia, sophisticated algorithmic tools are used to assess the quality of edits and take corrective actions. However, algorithms can fail to solve the problems they were designed for if they conflict with the values of communities who…

人机交互 · 计算机科学 2020-01-15 C. Estelle Smith , Bowen Yu , Anjali Srivastava , Aaron Halfaker , Loren Terveen , Haiyi Zhu

In this paper, we study the network of global interconnections between language communities, based on shared co-editing interests of Wikipedia editors, and show that although English is discussed as a potential lingua franca of the digital…

物理与社会 · 物理学 2016-03-15 Anna Samoilenko , Fariba karimi , Daniel Edler , Jérôme Kunegis , Markus Strohmaier

Wikipedia is a critical source of information for millions of users across the Web. It serves as a key resource for large language models, search engines, question-answering systems, and other Web-based applications. In Wikipedia, content…

Processing low-resource languages, such as Kiswahili, using machine learning is difficult due to lack of adequate training data. However, such low-resource languages are still important for human communication and are already in daily use…

计算与语言 · 计算机科学 2025-01-17 Barack Wamkaya Wanjawa , Lawrence Muchemi , Evans Miriti

ASR has achieved remarkable global progress, yet African low-resource languages remain rigorously underrepresented, producing barriers to digital inclusion across the continent with more than +2000 languages. This systematic literature…

Wikipedia plays a crucial role in the integrity of the Web. This work analyzes the reliability of this global encyclopedia through the lens of its references. We operationalize the notion of reference quality by defining reference need…

Peer production platforms like Wikipedia commonly suffer from content gaps. Prior research suggests recommender systems can help solve this problem, by guiding editors towards underrepresented topics. However, it remains unclear whether…

计算机与社会 · 计算机科学 2024-04-11 Mo Houtti , Isaac Johnson , Morten Warncke-Wang , Loren Terveen

Automatic Speech Recognition (ASR) technologies have transformed human-computer interaction; however, low-resource languages in Africa remain significantly underrepresented in both research and practical applications. This study…

Low-resource languages serve as invaluable repositories of human history, preserving cultural and intellectual diversity. Despite their significance, they remain largely absent from modern natural language processing systems. While progress…

计算与语言 · 计算机科学 2026-03-17 Offiong Bassey Edet , Mbuotidem Sunday Awak , Emmanuel Oyo-Ita , Benjamin Okon Nyong , Ita Etim Bassey

This study investigates the potential of Large Language Models (LLMs), particularly GPT-4o, for Optical Character Recognition (OCR) in low-resource scripts such as Urdu, Albanian, and Tajik, with English serving as a benchmark. Using a…

机器学习 · 计算机科学 2024-12-23 Muhammad Abdullah Sohail , Salaar Masood , Hamza Iqbal

The Internet facilitates large-scale collaborative projects and the emergence of Web 2.0 platforms, where producers and consumers of content unify, has drastically changed the information market. On the one hand, the promise of the "wisdom…

计算与语言 · 计算机科学 2023-01-05 Dong Nguyen , Barbara McGillivray , Taha Yasseri

We aim to investigate the performance of current OCR systems on low resource languages and low resource scripts. We introduce and make publicly available a novel benchmark, OCR4MT, consisting of real and synthetic data, enriched with noise,…

计算与语言 · 计算机科学 2022-03-15 Oana Ignat , Jean Maillard , Vishrav Chaudhary , Francisco Guzmán

Large Language Models (LLMs) like GPT-4 and LLaMA have shown incredible proficiency at natural language processing tasks and have even begun to excel at tasks across other modalities such as vision and audio. Despite their success, LLMs…

计算与语言 · 计算机科学 2024-03-12 Michael Andersland

Wikipedia is one of the main repositories of free knowledge available today, with a central role in the Web ecosystem. For this reason, it can also be a battleground for actors trying to impose specific points of view or even spreading…

计算机与社会 · 计算机科学 2021-07-01 Pablo Aragón , Diego Sáez-Trumper