中文
相关论文

相关论文: AfroLID: A Neural Language Identification Tool for…

200 篇论文

Large language models (LLMs) are increasingly multilingual, yet open models continue to underperform relative to proprietary systems, with the gap most pronounced for African languages. Continued pre-training (CPT) offers a practical route…

计算与语言 · 计算机科学 2026-05-06 Hao Yu , Tianyi Xu , Michael A. Hedderich , Wassim Hamidouche , Syed Waqas Zamir , David Ifeoluwa Adelani

Text classifiers are applied at scale in the form of one-size-fits-all solutions. Nevertheless, many studies show that classifiers are biased regarding different languages and dialects. When measuring and discovering these biases, some gaps…

计算与语言 · 计算机科学 2022-09-16 Brandon Lwowski , Paul Rad , Anthony Rios

Digital news platforms use news recommenders as the main instrument to cater to the individual information needs of readers. Despite an increasingly language-diverse online community, in which many Internet users consume news in multiple…

信息检索 · 计算机科学 2024-03-27 Andreea Iana , Goran Glavaš , Heiko Paulheim

This paper develops an approach to language identification in which the set of languages considered by the model depends on the geographic origin of the text in question. Given that many digital corpora can be geo-referenced at the country…

计算与语言 · 计算机科学 2024-03-18 Jonathan Dunn , Lane Edwards-Brown

Due to their crucial role in all NLP, several benchmarks have been proposed to evaluate pretrained language models. In spite of these efforts, no public benchmark of diverse nature currently exists for evaluation of Arabic. This makes it…

计算与语言 · 计算机科学 2023-05-31 AbdelRahim Elmadany , El Moatez Billah Nagoudi , Muhammad Abdul-Mageed

Natural language processing for the Turkic language family, spoken by over 200 million people across Eurasia, remains fragmented, with most languages lacking unified tooling and resources. We present TurkicNLP, an open-source Python library…

计算与语言 · 计算机科学 2026-05-25 Sherzod Hakimov

Due to a drastic improvement in the quality of internet services worldwide, there is an explosion of multilingual content generation and consumption. This is especially prevalent in countries with large multilingual audience, who are…

音频与语音处理 · 电气工程与系统科学 2020-07-14 Mudit Verma , Arun Balaji Buduru

Arabic has diverse dialects, where one dialect can be substantially different from the others. In the NLP literature, some assumptions about these dialects are widely adopted (e.g., ``Arabic dialects can be grouped into distinguishable…

计算与语言 · 计算机科学 2025-05-29 Amr Keleg , Sharon Goldwater , Walid Magdy

We describe findings of the third Nuanced Arabic Dialect Identification Shared Task (NADI 2022). NADI aims at advancing state of the art Arabic NLP, including on Arabic dialects. It does so by affording diverse datasets and modeling…

计算与语言 · 计算机科学 2022-10-24 Muhammad Abdul-Mageed , Chiyu Zhang , AbdelRahim Elmadany , Houda Bouamor , Nizar Habash

As language and speech technologies become more advanced, the lack of fundamental digital resources for African languages, such as data, spell checkers and Part of Speech taggers, means that the digital divide between these languages and…

计算与语言 · 计算机科学 2020-07-24 Kathleen Siminyu , Sackey Freshia , Jade Abbott , Vukosi Marivate

Accurately classifying accents and assessing accentedness in non-native speakers are both challenging tasks due to the complexity and diversity of accent and dialect variations. In this study, embeddings from advanced pre-trained language…

音频与语音处理 · 电气工程与系统科学 2023-10-18 Shahram Ghorbani , John H. L. Hansen

Africa has over 2000 indigenous languages but they are under-represented in NLP research due to lack of datasets. In recent years, there have been progress in developing labeled corpora for African languages. However, they are often…

计算与语言 · 计算机科学 2023-08-23 Iyanuoluwa Shode , David Ifeoluwa Adelani , Jing Peng , Anna Feldman

Multilingual search can be achieved with subword tokenization. The accuracy of traditional TF-IDF approaches depend on manually curated tokenization, stop words and stemming rules, whereas subword TF-IDF (STF-IDF) can offer higher accuracy…

计算与语言 · 计算机科学 2022-09-30 Artit Wangperawong

The video-language (VL) pretraining has achieved remarkable improvement in multiple downstream tasks. However, the current VL pretraining framework is hard to extend to multiple modalities (N modalities, N>=3) beyond vision and language. We…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Bin Zhu , Bin Lin , Munan Ning , Yang Yan , Jiaxi Cui , HongFa Wang , Yatian Pang , Wenhao Jiang , Junwu Zhang , Zongwei Li , Wancai Zhang , Zhifeng Li , Wei Liu , Li Yuan

Evaluating Large Language Models (LLMs) in low-resource and linguistically diverse languages remains a significant challenge in NLP, particularly for languages using non-Latin scripts like those spoken in India. Existing benchmarks…

计算与语言 · 计算机科学 2025-02-05 Sshubam Verma , Mohammed Safi Ur Rahman Khan , Vishwajeet Kumar , Rudra Murthy , Jaydeep Sen

Natural language inference (NLI) is known as one of the central tasks in natural language processing (NLP) which encapsulates many fundamental aspects of language understanding. With the considerable achievements of data-hungry deep…

Lombard, an underresourced language variety spoken by approximately 3.8 million people in Northern Italy and Southern Switzerland, lacks a unified orthographic standard. Multiple orthographic systems exist, creating challenges for NLP…

计算与语言 · 计算机科学 2026-03-31 Edoardo Signoroni , Pavel Rychlý

Today, the exponential rise of large models developed by academic and industrial institutions with the help of massive computing resources raises the question of whether someone without access to such resources can make a valuable…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Alexander Visheratin