中文
相关论文

相关论文: The WiLI benchmark dataset for written language id…

200 篇论文

Large-scale datasets have successively proven their fundamental importance in several research fields, especially for early progress in some emerging topics. In this paper, we focus on the problem of visual speech recognition, also known as…

计算机视觉与模式识别 · 计算机科学 2019-04-25 Shuang Yang , Yuanhang Zhang , Dalu Feng , Mingmin Yang , Chenhao Wang , Jingyun Xiao , Keyu Long , Shiguang Shan , Xilin Chen

Complex Word Identification (CWI) is the task of identifying which words or phrases in a sentence are difficult to understand by a target audience. The latest CWI Shared Task released data for two settings: monolingual (i.e. train and test…

Human label variation, or annotation disagreement, exists in many natural language processing (NLP) tasks, including natural language inference (NLI). To gain direct evidence of how NLI label variation arises, we build LiveNLI, an English…

计算与语言 · 计算机科学 2023-10-24 Nan-Jiang Jiang , Chenhao Tan , Marie-Catherine de Marneffe

We propose a new formulation for multilingual entity linking, where language-specific mentions resolve to a language-agnostic Knowledge Base. We train a dual encoder in this new setting, building on prior work with improved feature…

计算与语言 · 计算机科学 2020-11-06 Jan A. Botha , Zifei Shan , Daniel Gillick

In the age of large transformer language models, linguistic evaluation play an important role in diagnosing models' abilities and limitations on natural language understanding. However, current evaluation methods show some significant…

计算与语言 · 计算机科学 2022-05-05 Zeming Chen , Qiyue Gao

The rapid development of multilingual large language models (LLMs) highlights the need for high-quality, diverse, and well-curated multilingual datasets. In this paper, we introduce DCAD-2000 (Data Cleaning as Anomaly Detection), a…

计算与语言 · 计算机科学 2025-10-27 Yingli Shen , Wen Lai , Shuo Wang , Xueren Zhang , Kangyang Luo , Alexander Fraser , Maosong Sun

Natural Language Inference (NLI) is considered a representative task to test natural language understanding (NLU). In this work, we propose an extensible framework to collectively yet categorically test diverse Logical reasoning…

人工智能 · 计算机科学 2023-09-06 Ishan Tarunesh , Somak Aditya , Monojit Choudhury

The language identification task is a crucial fundamental step in NLP. Often it serves as a pre-processing step for widely used NLP applications such as multilingual machine translation, information retrieval, question and answering, and…

计算与语言 · 计算机科学 2026-01-08 Yash Ingle , Pruthwik Mishra

While Indic NLP has made rapid advances recently in terms of the availability of corpora and pre-trained models, benchmark datasets on standard NLU tasks are limited. To this end, we introduce IndicXNLI, an NLI dataset for 11 Indic…

计算与语言 · 计算机科学 2022-04-20 Divyanshu Aggarwal , Vivek Gupta , Anoop Kunchukuttan

Complex Word Identification (CWI) is a task centered on detecting hard-to-understand words, or groups of words, in texts from different areas of expertise. The purpose of CWI is to highlight problematic structures that non-native speakers…

计算与语言 · 计算机科学 2020-10-05 George-Eduard Zaharia , Dumitru-Clementin Cercel , Mihai Dascalu

Language Identification (LID) is a core task in multilingual NLP, yet current systems often overfit to clean, monolingual data. This work introduces DIVERS-BENCH, a comprehensive evaluation of state-of-the-art LID models across diverse…

计算与语言 · 计算机科学 2025-09-23 Jessica Ojo , Zina Kamel , David Ifeoluwa Adelani

We present a scalable, modular pipeline for automatic neologism detection that combines rule-based filtering with LLM classification. The pipeline is grounded in two complementary word-formation frameworks, grammatical and extra-grammatical…

计算与语言 · 计算机科学 2026-05-08 Diego Rossini , Lonneke van der Plas

Knowledge-based authentication is crucial for task-oriented spoken dialogue systems that offer personalised and privacy-focused services. Such systems should be able to enrol (E), verify (V), and identify (I) new and recurring users based…

计算与语言 · 计算机科学 2022-04-29 Georgios P. Spithourakis , Ivan Vulić , Michał Lis , Iñigo Casanueva , Paweł Budzianowski

Large Language Models (LLMs) are trained on vast amounts of data, most of which is automatically scraped from the internet. This data includes encyclopedic documents that harbor a vast amount of general knowledge (e.g., Wikipedia) but also…

The task of determining a speaker's native language based only on his speeches in a second language is known as Native Language Identification or NLI. Due to its increasing applications in various domains of speech signal processing, this…

计算与语言 · 计算机科学 2018-11-15 Ahmed Nazim Uddin , Md Ashequr Rahman , Md. Rafidul Islam , Mohammad Ariful Haque

Multilingual NLP often relies on dataset counts from centralized catalogues to characterize which languages are resource-rich or resource-poor. However, these catalogues record only one layer of dataset visibility: what has been registered…

计算与语言 · 计算机科学 2026-05-19 Zhiyin Tan , Changxu Duan

This paper addresses spoken language identification (SLI) and speech recognition of multilingual broadcast and institutional speech, real application scenarios that have been rarely addressed in the SLI literature. Observing that in these…

音频与语音处理 · 电气工程与系统科学 2024-06-14 Martina Valente , Fabio Brugnara , Giovanni Morrone , Enrico Zovato , Leonardo Badino

In this paper we present the Wikipedia Cultural Diversity dataset. For each existing Wikipedia language edition, the dataset contains a classification of the articles that represent its associated cultural context, i.e. all concepts and…

计算机与社会 · 计算机科学 2019-06-11 Marc Miquel-Ribé , David Laniado

The modern technological landscape has trended towards increased precision and greater digitization of information. However, the methods used to record and communicate scientific procedures have remained largely unchanged over the last…

数字图书馆 · 计算机科学 2024-09-11 Robert W. Epps , Amanda A. Volk , Robert R. White , Robert Tirawat , Rosemary C. Bramante , Joseph J. Berry

Existing datasets for natural language inference (NLI) have propelled research on language understanding. We propose a new method for automatically deriving NLI datasets from the growing abundance of large-scale question answering datasets.…

计算与语言 · 计算机科学 2018-09-12 Dorottya Demszky , Kelvin Guu , Percy Liang