中文
相关论文

相关论文: Mining Wikidata for Name Resources for African Lan…

200 篇论文

Creating speech datasets, models, and evaluation frameworks for low-resource languages remains challenging given the lack of a broad base of pertinent experience to draw from. This paper reports on the field collection of 612 hours of…

Researchers working on low-resource languages face persistent challenges due to limited data availability and restricted access to computational resources. Although most large language models (LLMs) are predominantly trained in…

计算与语言 · 计算机科学 2025-05-27 Odunayo Ogundepo , Akintunde Oladipo , Kelechi Ogueji , Esther Adenuga , David Ifeoluwa Adelani , Jimmy Lin

Wikipedia is a critical source of information for millions of users across the Web. It serves as a key resource for large language models, search engines, question-answering systems, and other Web-based applications. In Wikipedia, content…

Wikimedia content is used extensively by the AI community and within the language modeling community in particular. In this paper, we provide a review of the different ways in which Wikimedia data is curated to use in NLP tasks across…

计算机与社会 · 计算机科学 2024-10-14 Isaac Johnson , Lucie-Aimée Kaffee , Miriam Redi

Large Language Models (LLMs) like GPT-4 and LLaMA have shown incredible proficiency at natural language processing tasks and have even begun to excel at tasks across other modalities such as vision and audio. Despite their success, LLMs…

计算与语言 · 计算机科学 2024-03-12 Michael Andersland

Consistency is one of the keys to maintainable source code and hence a successful software project. We propose a novel method of extracting the intent of programmers from source code of a large project (~300kLOC) and checking the semantic…

软件工程 · 计算机科学 2022-03-21 Yusuke Shinyama , Yoshitaka Arahori , Katsuhiko Gondow

Datasets for data-to-text generation typically focus either on multi-domain, single-sentence generation or on single-domain, long-form generation. In this work, we cast generating Wikipedia sections as a data-to-text generation task and…

计算与语言 · 计算机科学 2021-06-03 Mingda Chen , Sam Wiseman , Kevin Gimpel

This paper addresses the harmonization of metadata from diverse repositories of language resources (LRs). Leveraging linked data and RDF techniques, we integrate data from multiple sources into a unified model based on DCAT and META-SHARE…

计算与语言 · 计算机科学 2025-01-13 Zixuan Liang

Wikidata and Wikipedia have been proven useful for reason-ing in natural language applications, like question answering or entitylinking. Yet, no existing work has studied the potential of Wikidata for commonsense reasoning. This paper…

人工智能 · 计算机科学 2020-10-19 Filip Ilievski , Pedro Szekely , Daniel Schwabe

Introduction: Scholarly research spans multiple languages, making multilingual metadata crucial for organizing and accessing knowledge across linguistic boundaries. These multilingual metadata already exist and are propagated throughout…

数字图书馆 · 计算机科学 2025-03-18 Dennis Donathan , Mike Nason , Marco Tullney , Julie Shi , Juan Pablo Alperin

Wikipedia serves as a globally accessible knowledge source with content in over 300 languages. Despite covering the same topics, the different versions of Wikipedia are written and updated independently. This leads to factual…

计算与语言 · 计算机科学 2026-05-19 Silvia Cappa , Lingxiao Kong , Pille-Riin Peet , Fanfu Wei , Yuchen Zhou , Jan-Christoph Kalo

In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The…

计算与语言 · 计算机科学 2021-03-11 Israel Abebe Azime , Nebil Mohammed

Processing low-resource languages, such as Kiswahili, using machine learning is difficult due to lack of adequate training data. However, such low-resource languages are still important for human communication and are already in daily use…

计算与语言 · 计算机科学 2025-01-17 Barack Wamkaya Wanjawa , Lawrence Muchemi , Evans Miriti

Large language models show human-like performance in knowledge extraction, reasoning and dialogue, but it remains controversial whether this performance is best explained by memorization and pattern matching, or whether it reflects…

机器学习 · 计算机科学 2023-08-30 Mathias Lykke Gammelgaard , Jonathan Gabel Christiansen , Anders Søgaard

We propose a new formulation for multilingual entity linking, where language-specific mentions resolve to a language-agnostic Knowledge Base. We train a dual encoder in this new setting, building on prior work with improved feature…

计算与语言 · 计算机科学 2020-11-06 Jan A. Botha , Zifei Shan , Daniel Gillick

The ability to have the same experience for different user groups (i.e., accessibility) is one of the most important characteristics of Web-based systems. The same is true for Knowledge Graph Question Answering (KGQA) systems that provide…

计算与语言 · 计算机科学 2022-02-08 Aleksandr Perevalov , Dennis Diefenbach , Ricardo Usbeck , Andreas Both

Natural Language Processing is a crucial frontier in artificial intelligence, with broad applications in many areas, including public health, agriculture, education, and commerce. However, due to the lack of substantial linguistic…

计算与语言 · 计算机科学 2025-01-22 Audrey Mbogho , Quin Awuor , Andrew Kipkebut , Lilian Wanzare , Vivian Oloo

Emotion lexicons describe the affective meaning of words and thus constitute a centerpiece for advanced sentiment and emotion analysis. Yet, manually curated lexicons are only available for a handful of languages, leaving most languages of…

计算与语言 · 计算机科学 2020-05-13 Sven Buechel , Susanna Rücker , Udo Hahn

We provide the largest compiled publicly available dictionaries of first, middle, and last names for the purpose of imputing race and ethnicity using, for example, Bayesian Improved Surname Geocoding (BISG). The dictionaries are based on…

其他统计学 · 统计学 2022-08-29 Evan T. R. Rosenman , Santiago Olivella , Kosuke Imai

The Linked Open Data practice has led to a significant growth of structured data on the Web in the last decade. Such structured data describe real-world entities in a machine-readable way, and have created an unprecedented opportunity for…

计算与语言 · 计算机科学 2022-06-27 Ziqi Zhang , Xingyi Song