中文
相关论文

相关论文: Kr\'eyoLID From Language Identification Towards La…

200 篇论文

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our methodology for mining such data from…

计算与语言 · 计算机科学 2015-09-30 Krzysztof Wołk , Krzysztof Marasek

Automatic terminology processing appeared 10 years ago when electronic corpora became widely available. Such processing may be statistically or linguistically based and produces terminology resources that can be used in a number of…

计算机与社会 · 计算机科学 2014-12-16 C. Enguehard , B. Daille , E. Morin

Whether or not several Creole languages which developed during the early modern period can be considered genetic descendants of European languages has been the subject of intense debate. This is in large part due to the absence of evidence…

计算与语言 · 计算机科学 2024-08-09 Rasul Dent , Juliette Janès , Thibault Clérice , Pedro Ortiz Suarez , Benoît Sagot

The rapid development of multilingual large language models (LLMs) highlights the need for high-quality, diverse, and well-curated multilingual datasets. In this paper, we introduce DCAD-2000 (Data Cleaning as Anomaly Detection), a…

计算与语言 · 计算机科学 2025-10-27 Yingli Shen , Wen Lai , Shuo Wang , Xueren Zhang , Kangyang Luo , Alexander Fraser , Maosong Sun

The dissemination of Large Language Models (LLMs), trained at scale, and endowed with powerful text-generating abilities, has made it easier for all to produce harmful, toxic, faked or forged content. In response, various proposals have…

计算与语言 · 计算机科学 2025-06-12 Matthieu Dubois , François Yvon , Pablo Piantanida

Language detoxification involves removing toxicity from offensive language. While a neutral-toxic paired dataset provides a straightforward approach for training detoxification models, creating such datasets presents several challenges: i)…

计算与语言 · 计算机科学 2025-06-17 Minkyeong Jeon , Hyemin Jeong , Yerang Kim , Jiyoung Kim , Jae Hyeon Cho , Byung-Jun Lee

In this paper, I describe several approaches to automatic or semi-automatic development of symbolic rules for grammar checkers from the information contained in corpora. The rules obtained this way are an important addition to…

计算与语言 · 计算机科学 2012-11-30 Marcin Miłkowski

Lexical resources are crucial for cross-linguistic analysis and can provide new insights into computational models for natural language learning. Here, we present an advanced database for comparative studies of words with multiple meanings,…

计算与语言 · 计算机科学 2025-08-22 Annika Tjuka , Robert Forkel , Christoph Rzymski , Johann-Mattis List

The performance of large language models (LLMs) is deeply influenced by the quality and composition of their training data. While much of the existing work has centered on English, there remains a gap in understanding how to construct…

计算与语言 · 计算机科学 2025-09-11 Thales Sales Almeida , Rodrigo Nogueira , Helio Pedrini

Available corpora for Argument Mining differ along several axes, and one of the key differences is the presence (or absence) of discourse markers to signal argumentative content. Exploring effective ways to use discourse markers has…

计算与语言 · 计算机科学 2023-06-08 Gil Rocha , Henrique Lopes Cardoso , Jonas Belouadi , Steffen Eger

The task of discovering topics in text corpora has been dominated by Latent Dirichlet Allocation and other Topic Models for over a decade. In order to apply these approaches to massive text corpora, the vocabulary needs to be reduced…

计算与语言 · 计算机科学 2019-08-08 Gibran Fuentes-Pineda , Ivan Vladimir Meza-Ruiz

Pre-training text representations have led to significant improvements in many areas of natural language processing. The quality of these models benefits greatly from the size of the pretraining corpora as long as its quality is preserved.…

Spoken language identification refers to the task of automatically predicting the spoken language in a given utterance. Conventionally, it is modeled as a speech-based language identification task. Prior techniques have been constrained to…

When looking at the structure of natural language, "phrases" and "words" are central notions. We consider the problem of identifying such "meaningful subparts" of language of any length and underlying composition principles in a completely…

计算与语言 · 计算机科学 2016-02-19 Stefan Gerdjikov , Klaus U. Schulz

This paper discusses creating and analysing a new dataset for data mining and text analytics research, contributing to a joint Leeds University research project for the Corpus of National Dialects. This report investigates machine learning…

计算与语言 · 计算机科学 2022-08-02 Omar Shaur Choudhry , Paul Omara Odida , Joshua Reiner , Keiron Appleyard , Danielle Kushnir , William Toon

Cross-lingual information retrieval (CLIR) addresses the challenge of retrieving relevant documents written in languages different from that of the original query. Research in this area has typically framed the task as monolingual retrieval…

信息检索 · 计算机科学 2025-10-02 Roksana Goworek , Olivia Macmillan-Scott , Eda B. Özyiğit

Multilingual text processing is useful because the information content found in different languages is complementary, both regarding facts and opinions. While Information Extraction and other text mining software can, in principle, be…

计算与语言 · 计算机科学 2014-01-14 Ralf Steinberger

A majority of language technologies are tailored for a small number of high-resource languages, while relatively many low-resource languages are neglected. One such group, Creole languages, have long been marginalized in academic study,…

We present an analysis pipeline and best practice guidelines for building and curating corpora of everyday conversation in diverse languages. Surveying language documentation corpora and other resources that cover 67 languages and varieties…

计算与语言 · 计算机科学 2022-05-11 Andreas Liesenfeld , Mark Dingemanse

Language development experts need tools that can automatically identify languages from fluent, conversational speech, and provide reliable estimates of usage rates at the level of an individual recording. However, language identification…

音频与语音处理 · 电气工程与系统科学 2023-05-31 Suzy J. Styles , Victoria Y. H. Chua , Fei Ting Woon , Hexin Liu , Leibny Paola Garcia Perera , Sanjeev Khudanpur , Andy W. H. Khong , Justin Dauwels