中文
相关论文

相关论文: Low-resourced Languages and Online Knowledge Repos…

200 篇论文

In this paper, we analyze the capabilities of the multi-lingual Dense Passage Retriever (mDPR) for extremely low-resource languages. In the Cross-lingual Open-Retrieval Answer Generation (CORA) pipeline, mDPR achieves success on…

信息检索 · 计算机科学 2024-08-23 Jie Wu , Zhaochun Ren , Suzan Verberne

The success of several architectures to learn semantic representations from unannotated text and the availability of these kind of texts in online multilingual resources such as Wikipedia has facilitated the massive and automatic creation…

计算与语言 · 计算机科学 2020-03-31 Jesujoba O. Alabi , Kwabena Amponsah-Kaakyire , David I. Adelani , Cristina España-Bonet

Question answering (QA) is the task of answering questions posed in natural language with free-form natural language answers extracted from a given passage. In the OpenQA variant, only a question text is given, and the system must retrieve…

计算与语言 · 计算机科学 2024-11-21 Emrah Budur , Rıza Özçelik , Dilara Soylu , Omar Khattab , Tunga Güngör , Christopher Potts

In recent decades, there has been a major shift towards improved digital access to scholarly works. However, even now that these works are available in digital form, they remain document-based, making it difficult to communicate the…

数字图书馆 · 计算机科学 2021-08-12 Oliver Karras , Eduard C. Groen , Javed Ali Khan , Sören Auer

Over the last few years, verifying the credibility of information sources has become a fundamental need to combat disinformation. Here, we present a language-agnostic model designed to assess the reliability of web domains as sources in…

社会与信息网络 · 计算机科学 2025-11-21 Jacopo D'Ignazi , Andreas Kaltenbrunner , Yelena Mejova , Michele Tizzani , Kyriaki Kalimeri , Mariano Beiró , Pablo Aragón

Online knowledge communities (OKC) such as Stack Exchange, Reddit, and Zhihu have long functioned as socio technical infrastructures for collective problem solving. The rapid adoption of Generative AI (GenAI) introduces both complementarity…

人机交互 · 计算机科学 2026-03-31 Ching Christie Pang , Xuetong Wang , Yuk Hang Tsui , Pan Hui

Lexical-semantic resources (LSRs), such as online lexicons and wordnets, are fundamental to natural language processing applications as well as to fields such as linguistic anthropology and language preservation. In many languages, however,…

计算与语言 · 计算机科学 2025-11-21 Hadi Khalilia , Jahna Otterbacher , Gabor Bella , Shandy Darma , Fausto Giunchiglia

Most social media users come from the Global South, where harmful content usually appears in local languages. Yet, AI-driven moderation systems struggle with low-resource languages spoken in these regions. Through semi-structured interviews…

计算与语言 · 计算机科学 2025-08-06 Farhana Shahid , Mona Elswah , Aditya Vashistha

Faced with a considerable lack of resources in African languages to carry out work in Natural Language Processing (NLP), Natural Language Understanding (NLU) and artificial intelligence, the research teams of NTeALan association has set…

计算与语言 · 计算机科学 2021-04-01 Elvis Mboning Tchiaze

Lack of encyclopedic text contributors, especially on Wikipedia, makes automated text generation for low resource (LR) languages a critical problem. Existing work on Wikipedia text generation has focused on English only where English…

计算与语言 · 计算机科学 2023-05-17 Dhaval Taunk , Shivprasad Sagare , Anupam Patil , Shivansh Subramanian , Manish Gupta , Vasudeva Varma

This survey delves into the current state of natural language processing (NLP) for four Ethiopian languages: Amharic, Afaan Oromo, Tigrinya, and Wolaytta. Through this paper, we identify key challenges and opportunities for NLP research in…

There are over a billion websites on the Internet that can potentially serve as sources of information on various topics. One of the most popular examples of such an online source is Wikipedia. This public knowledge base is co-edited by…

信息检索 · 计算机科学 2022-05-02 Włodzimierz Lewoniewski , Krzysztof Węcel , Witold Abramowicz

Online platforms, particularly Wikipedia, have become critical infrastructures for providing diverse linguistic and cultural contexts. This human-curated knowledge now forms the foundation for modern AI. However, we have not yet fully…

计算机与社会 · 计算机科学 2025-07-31 Akira Matsui , Fujio Toriumi , Mitsuo Yoshida , Taichi Murayama , Shiori Hironaka

In addition to user-generated content, Open Educational Resources are increasingly made available on the Web by several institutions and organizations with the aim of being re-used. Nevertheless, it is still difficult for users to find…

人机交互 · 计算机科学 2015-09-10 Jaspreet Singh , Zeon Trevor Fernando , Saniya Chawla

Large Language Models (LLMs) have demonstrated remarkable success across a wide range of tasks and domains. However, their performance in low-resource language translation, particularly when translating into these languages, remains…

The use of domain knowledge is generally found to improve query efficiency in content filtering applications. In particular, tangible benefits have been achieved when using knowledge-based approaches within more specialized fields, such as…

信息检索 · 计算机科学 2015-03-17 Pekka Malo , Pyry Siitari , Oskar Ahlgren , Jyrki Wallenius , Pekka Korhonen

The usage of non-authoritative data for disaster management presents the opportunity of accessing timely information that might not be available through other means, as well as the challenge of dealing with several layers of biases.…

信息检索 · 计算机科学 2020-01-27 Valerio Lorini , Javier Rando , Diego Saez-Trumper , Carlos Castillo

To effectively navigate the AI revolution, AI literacy is crucial. However, content predominantly exists in dominant languages, creating a gap for low-resource languages like Yoruba (41 million native speakers). This case study explores…

计算与语言 · 计算机科学 2024-03-11 Wuraola Oyewusi

Many natural language processing (NLP) tasks make use of massively pre-trained language models, which are computationally expensive. However, access to high computational resources added to the issue of data scarcity of African languages…

With 60M articles in more than 300 language versions, Wikipedia is the largest platform for open and freely accessible knowledge. While the available content has been growing continuously at a rate of around 200K new articles each month,…

社会与信息网络 · 计算机科学 2024-10-08 Akhil Arora , Robert West , Martin Gerlach