English
Related papers

Related papers: Low-resourced Languages and Online Knowledge Repos…

200 papers

In this paper, we analyze the capabilities of the multi-lingual Dense Passage Retriever (mDPR) for extremely low-resource languages. In the Cross-lingual Open-Retrieval Answer Generation (CORA) pipeline, mDPR achieves success on…

Information Retrieval · Computer Science 2024-08-23 Jie Wu , Zhaochun Ren , Suzan Verberne

The success of several architectures to learn semantic representations from unannotated text and the availability of these kind of texts in online multilingual resources such as Wikipedia has facilitated the massive and automatic creation…

Computation and Language · Computer Science 2020-03-31 Jesujoba O. Alabi , Kwabena Amponsah-Kaakyire , David I. Adelani , Cristina España-Bonet

Question answering (QA) is the task of answering questions posed in natural language with free-form natural language answers extracted from a given passage. In the OpenQA variant, only a question text is given, and the system must retrieve…

Computation and Language · Computer Science 2024-11-21 Emrah Budur , Rıza Özçelik , Dilara Soylu , Omar Khattab , Tunga Güngör , Christopher Potts

In recent decades, there has been a major shift towards improved digital access to scholarly works. However, even now that these works are available in digital form, they remain document-based, making it difficult to communicate the…

Digital Libraries · Computer Science 2021-08-12 Oliver Karras , Eduard C. Groen , Javed Ali Khan , Sören Auer

Over the last few years, verifying the credibility of information sources has become a fundamental need to combat disinformation. Here, we present a language-agnostic model designed to assess the reliability of web domains as sources in…

Social and Information Networks · Computer Science 2025-11-21 Jacopo D'Ignazi , Andreas Kaltenbrunner , Yelena Mejova , Michele Tizzani , Kyriaki Kalimeri , Mariano Beiró , Pablo Aragón

Online knowledge communities (OKC) such as Stack Exchange, Reddit, and Zhihu have long functioned as socio technical infrastructures for collective problem solving. The rapid adoption of Generative AI (GenAI) introduces both complementarity…

Human-Computer Interaction · Computer Science 2026-03-31 Ching Christie Pang , Xuetong Wang , Yuk Hang Tsui , Pan Hui

Lexical-semantic resources (LSRs), such as online lexicons and wordnets, are fundamental to natural language processing applications as well as to fields such as linguistic anthropology and language preservation. In many languages, however,…

Computation and Language · Computer Science 2025-11-21 Hadi Khalilia , Jahna Otterbacher , Gabor Bella , Shandy Darma , Fausto Giunchiglia

Most social media users come from the Global South, where harmful content usually appears in local languages. Yet, AI-driven moderation systems struggle with low-resource languages spoken in these regions. Through semi-structured interviews…

Computation and Language · Computer Science 2025-08-06 Farhana Shahid , Mona Elswah , Aditya Vashistha

Faced with a considerable lack of resources in African languages to carry out work in Natural Language Processing (NLP), Natural Language Understanding (NLU) and artificial intelligence, the research teams of NTeALan association has set…

Computation and Language · Computer Science 2021-04-01 Elvis Mboning Tchiaze

Lack of encyclopedic text contributors, especially on Wikipedia, makes automated text generation for low resource (LR) languages a critical problem. Existing work on Wikipedia text generation has focused on English only where English…

Computation and Language · Computer Science 2023-05-17 Dhaval Taunk , Shivprasad Sagare , Anupam Patil , Shivansh Subramanian , Manish Gupta , Vasudeva Varma

This survey delves into the current state of natural language processing (NLP) for four Ethiopian languages: Amharic, Afaan Oromo, Tigrinya, and Wolaytta. Through this paper, we identify key challenges and opportunities for NLP research in…

There are over a billion websites on the Internet that can potentially serve as sources of information on various topics. One of the most popular examples of such an online source is Wikipedia. This public knowledge base is co-edited by…

Information Retrieval · Computer Science 2022-05-02 Włodzimierz Lewoniewski , Krzysztof Węcel , Witold Abramowicz

Online platforms, particularly Wikipedia, have become critical infrastructures for providing diverse linguistic and cultural contexts. This human-curated knowledge now forms the foundation for modern AI. However, we have not yet fully…

Computers and Society · Computer Science 2025-07-31 Akira Matsui , Fujio Toriumi , Mitsuo Yoshida , Taichi Murayama , Shiori Hironaka

In addition to user-generated content, Open Educational Resources are increasingly made available on the Web by several institutions and organizations with the aim of being re-used. Nevertheless, it is still difficult for users to find…

Human-Computer Interaction · Computer Science 2015-09-10 Jaspreet Singh , Zeon Trevor Fernando , Saniya Chawla

Large Language Models (LLMs) have demonstrated remarkable success across a wide range of tasks and domains. However, their performance in low-resource language translation, particularly when translating into these languages, remains…

The use of domain knowledge is generally found to improve query efficiency in content filtering applications. In particular, tangible benefits have been achieved when using knowledge-based approaches within more specialized fields, such as…

Information Retrieval · Computer Science 2015-03-17 Pekka Malo , Pyry Siitari , Oskar Ahlgren , Jyrki Wallenius , Pekka Korhonen

The usage of non-authoritative data for disaster management presents the opportunity of accessing timely information that might not be available through other means, as well as the challenge of dealing with several layers of biases.…

Information Retrieval · Computer Science 2020-01-27 Valerio Lorini , Javier Rando , Diego Saez-Trumper , Carlos Castillo

To effectively navigate the AI revolution, AI literacy is crucial. However, content predominantly exists in dominant languages, creating a gap for low-resource languages like Yoruba (41 million native speakers). This case study explores…

Computation and Language · Computer Science 2024-03-11 Wuraola Oyewusi

Many natural language processing (NLP) tasks make use of massively pre-trained language models, which are computationally expensive. However, access to high computational resources added to the issue of data scarcity of African languages…

With 60M articles in more than 300 language versions, Wikipedia is the largest platform for open and freely accessible knowledge. While the available content has been growing continuously at a rate of around 200K new articles each month,…

Social and Information Networks · Computer Science 2024-10-08 Akhil Arora , Robert West , Martin Gerlach