中文
相关论文

相关论文: Short Text Language Identification for Under Resou…

200 篇论文

Virtual assistants and text chatbots have recently been gaining popularity. Given the short message nature of text-based chat interactions, the language identification systems of these bots might only have 15 or 20 characters to make a…

计算与语言 · 计算机科学 2017-11-02 Bernardt Duvenhage , Mfundo Ntini , Phala Ramonyai

Language identification (LID) is a critical step in curating multilingual LLM pretraining corpora from web crawls. While many studies on LID model training focus on collecting diverse training data to improve performance, low-resource…

计算与语言 · 计算机科学 2026-03-11 Negar Foroutan , Jakhongir Saydaliev , Ye Eun Kim , Antoine Bosselut

Language Identification (LID) is a challenging task, especially when the input texts are short and noisy such as posts and statuses on social media or chat logs on gaming forums. The task has been tackled by either designing a feature set…

计算与语言 · 计算机科学 2019-10-16 Duy Tin Vo , Richard Khoury

With the constant growth of the World Wide Web and the number of documents in different languages accordingly, the need for reliable language detection tools has increased as well. Platforms such as Twitter with predominantly short texts…

计算与语言 · 计算机科学 2016-08-31 Ivana Balazevic , Mikio Braun , Klaus-Robert Müller

Informal transliteration from other languages to English is prevalent in social media threads, instant messaging, and discussion forums. Without identifying the language of such transliterated text, users who do not speak that language…

音频与语音处理 · 电气工程与系统科学 2020-04-29 Sourav Sen

Several recent papers have published good solutions for language identification (LID) for about 300 high-resource and medium-resource languages. However, there is no LID available that (i) covers a wide range of low-resource languages, (ii)…

计算与语言 · 计算机科学 2024-07-04 Amir Hossein Kargaran , Ayyoob Imani , François Yvon , Hinrich Schütze

The increase in the use of microblogging came along with the rapid growth on short linguistic data. On the other hand deep learning is considered to be the new frontier to extract meaningful information out of large amount of raw data in an…

计算与语言 · 计算机科学 2017-01-16 Priyank Mathur , Arkajyoti Misra , Emrah Budur

Language Identification (LID) is an important component of many multilingual natural language processing pipelines, where it facilitates corpus curation, training data analysis, and cross-lingual evaluation of large language models. Despite…

计算与语言 · 计算机科学 2026-02-20 Clara Meister , Ahmetcan Yavuz , Pietro Lesci , Tiago Pimentel

Language Identification is the task of identifying a document's language. For applications like automatic spell checker selection, language identification must use very short strings such as text message fragments. In this work, we…

计算与语言 · 计算机科学 2021-02-15 Mads Toftrup , Søren Asger Sørensen , Manuel R. Ciosici , Ira Assent

This memo describes NTR/TSU winning submission for Low Resource ASR challenge at Dialog2021 conference, language identification track. Spoken Language Identification (LID) is an important step in a multilingual Automated Speech Recognition…

音频与语音处理 · 电气工程与系统科学 2021-06-02 Roman Bedyakin , Nikolay Mikhaylovskiy

Language Identification (LID) systems are used to classify the spoken language from a given audio sample and are typically the first step for many spoken language processing tasks, such as Automatic Speech Recognition (ASR) systems. Without…

计算机视觉与模式识别 · 计算机科学 2017-08-17 Christian Bartz , Tom Herold , Haojin Yang , Christoph Meinel

Language Identification (LID) is the task of determining the language of a given text and is a fundamental preprocessing step that affects the reliability of downstream NLP applications. While recent work has expanded LID coverage for…

计算与语言 · 计算机科学 2026-01-30 Sang Yun Kwon , AbdelRahim Elmadany , Muhammad Abdul-Mageed

Local/Native South African languages are classified as low-resource languages. As such, it is essential to build the resources for these languages so that they can benefit from advances in the field of natural language processing. In this…

计算与语言 · 计算机科学 2023-06-14 Andani Madodonga , Vukosi Marivate , Matthew Adendorff

Spoken language identification (LID) technologies have improved in recent years from discriminating largely distinct languages to discriminating highly similar languages or even dialects of the same language. One aspect that has been mostly…

音频与语音处理 · 电气工程与系统科学 2020-01-30 Andrew Titus , Jan Silovsky , Nanxin Chen , Roger Hsiao , Mary Young , Arnab Ghoshal

Language models are the foundation of current neural network-based models for natural language understanding and generation. However, research on the intrinsic performance of language models on African languages has been extremely limited,…

计算与语言 · 计算机科学 2021-04-05 Stuart Mesham , Luc Hayward , Jared Shapiro , Jan Buys

Language-independent tokenisation (LIT) methods that do not require labelled language resources or lexicons have recently gained popularity because of their applicability in resource-poor languages. Moreover, they compactly represent a…

计算与语言 · 计算机科学 2020-02-26 Danushka Bollegala , Ryuichi Kiryo , Kosuke Tsujino , Haruki Yukawa

Recognizing semantically similar sentences or paragraphs across languages is beneficial for many tasks, ranging from cross-lingual information retrieval and plagiarism detection to machine translation. Recently proposed methods for…

计算与语言 · 计算机科学 2018-01-22 Goran Glavaš , Marc Franco-Salvador , Simone Paolo Ponzetto , Paolo Rosso

Language identification (LID) is a crucial precursor for NLP, especially for mining web data. Problematically, most of the world's 7000+ languages today are not covered by LID technologies. We address this pressing issue for Africa by…

计算与语言 · 计算机科学 2022-12-08 Ife Adebara , AbdelRahim Elmadany , Muhammad Abdul-Mageed , Alcides Alcoba Inciarte

Identifying closely related languages at sentence level is difficult, in particular because it is often impossible to assign a sentence to a single language. In this paper, we focus on multi-label sentence-level Scandinavian language…

We present MaskLID, a simple, yet effective, code-switching (CS) language identification (LID) method. MaskLID does not require any training and is designed to complement current high-performance sentence-level LIDs. Sentence-level LIDs are…

计算与语言 · 计算机科学 2024-06-11 Amir Hossein Kargaran , François Yvon , Hinrich Schütze
‹ 上一页 1 2 3 10 下一页 ›