中文
相关论文

相关论文: Mining Wikidata for Name Resources for African Lan…

200 篇论文

English Wikipedia has long been an important data source for much research and natural language machine learning modeling. The growth of non-English language editions of Wikipedia, greater computational resources, and calls for equity in…

计算机与社会 · 计算机科学 2022-04-07 Isaac Johnson , Emily Lescak

With the advent of open source software, a veritable treasure trove of previously proprietary software development data was made available. This opened the field of empirical software engineering research to anyone in academia. Data that is…

软件工程 · 计算机科学 2022-04-19 Adam Tutko , Austin Z. Henley , Audris Mockus

Progress in cross-lingual modeling depends on challenging, realistic, and diverse evaluation sets. We introduce Multilingual Knowledge Questions and Answers (MKQA), an open-domain question answering evaluation set comprising 10k…

计算与语言 · 计算机科学 2021-08-18 Shayne Longpre , Yi Lu , Joachim Daiber

Wikipedia is the largest online encyclopedia, used by algorithms and web users as a central hub of reliable information on the web. The quality and reliability of Wikipedia content is maintained by a community of volunteer editors. Machine…

信息检索 · 计算机科学 2021-06-02 KayYen Wong , Miriam Redi , Diego Saez-Trumper

The Speech Wikimedia Dataset is a publicly available compilation of audio with transcriptions extracted from Wikimedia Commons. It includes 1780 hours (195 GB) of CC-BY-SA licensed transcribed speech from a diverse set of scenarios and…

Deep neural language models such as BERT have enabled substantial recent advances in many natural language processing tasks. Due to the effort and computational cost involved in their pre-training, language-specific models are typically…

计算与语言 · 计算机科学 2020-06-03 Sampo Pyysalo , Jenna Kanerva , Antti Virtanen , Filip Ginter

Recent advancements in Natural Language Processing (NLP) has led to the proliferation of large pretrained language models. These models have been shown to yield good performance, using in-context learning, even on unseen tasks and…

计算与语言 · 计算机科学 2023-05-12 Jessica Ojo , Kelechi Ogueji

West African Pidgin English is a language that is significantly spoken in West Africa, consisting of at least 75 million speakers. Nevertheless, proper machine translation systems and relevant NLP datasets for pidgin English are virtually…

计算与语言 · 计算机科学 2021-04-28 Ernie Chang , David Ifeoluwa Adelani , Xiaoyu Shen , Vera Demberg

We address a notable gap in Natural Language Processing (NLP) by introducing a collection of resources designed to improve Machine Translation (MT) for low-resource languages, with a specific focus on African languages. First, we introduce…

计算与语言 · 计算机科学 2024-07-15 AbdelRahim Elmadany , Ife Adebara , Muhammad Abdul-Mageed

The Data Web refers to the vast and rapidly increasing quantity of scientific, corporate, government and crowd-sourced data published in the form of Linked Open Data, which encourages the uniform representation of heterogeneous data items…

Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such assumptions of high quality have become the subject of scrutiny in low-resource and…

The lack of comprehensive, high-quality health data in developing nations creates a roadblock for combating the impacts of disease. One key challenge is understanding the health information needs of people in these nations. Without…

计算机与社会 · 计算机科学 2019-04-18 Rediet Abebe , Shawndra Hill , Jennifer Wortman Vaughan , Peter M. Small , H. Andrew Schwartz

Confidently making progress on multilingual modeling requires challenging, trustworthy evaluations. We present TyDi QA---a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs. The languages…

Information presented in Wikipedia articles must be attributable to reliable published sources in the form of references. This study examines over 5 million Wikipedia articles to assess the reliability of references in multiple language…

计算机与社会 · 计算机科学 2023-09-06 Aitolkyn Baigutanova , Diego Saez-Trumper , Miriam Redi , Meeyoung Cha , Pablo Aragón

While the similarity between two concept words has been evaluated and studied for decades, much less attention has been devoted to algorithms that can compute the similarity of nodes in very large knowledge graphs, like Wikidata. To…

人工智能 · 计算机科学 2021-08-13 Filip Ilievski , Pedro Szekely , Gleb Satyukov , Amandeep Singh

Public knowledge graphs such as DBpedia and Wikidata have been recognized as interesting sources of background knowledge to build content-based recommender systems. They can be used to add information about the items to be recommended and…

信息检索 · 计算机科学 2021-05-04 Michael Matthias Voit , Heiko Paulheim

Wikipedia is a goldmine of information; not just for its many readers, but also for the growing community of researchers who recognize it as a resource of exceptional scale and utility. It represents a vast investment of manual effort and…

人工智能 · 计算机科学 2009-05-10 Olena Medelyan , David Milne , Catherine Legg , Ian H. Witten

As machine learning and data science applications grow ever more prevalent, there is an increased focus on data sharing and open data initiatives, particularly in the context of the African continent. Many argue that data sharing can…

计算机与社会 · 计算机科学 2021-03-02 Rediet Abebe , Kehinde Aruleba , Abeba Birhane , Sara Kingsley , George Obaido , Sekou L. Remy , Swathi Sadagopan

We present a Wikidata-based framework, called KIF, for virtually integrating heterogeneous knowledge sources. KIF is written in Python and is released as open-source. It leverages Wikidata's data model and vocabulary plus user-defined…

‹ 上一页 1 8 9 10 下一页 ›