中文
相关论文

相关论文: Improving Informally Romanized Language Identifica…

200 篇论文

Multilingual writers and speakers often alternate between two languages in a single discourse, a practice called "code-switching". Existing sentiment detection methods are usually trained on sentiment-labeled monolingual text. Manually…

计算与语言 · 计算机科学 2019-06-14 Bidisha Samanta , Niloy Ganguly , Soumen Chakrabarti

In the context of pretraining of Large Language Models (LLMs), synthetic data has emerged as an alternative for generating high-quality pretraining data at scale. This is particularly beneficial in low-resource language settings where the…

We present the largest publicly available synthetic OCR benchmark dataset for Indic languages. The collection contains a total of 90k images and their ground truth for 23 Indic languages. OCR model validation in Indic languages require a…

计算机视觉与模式识别 · 计算机科学 2022-05-06 Naresh Saini , Promodh Pinto , Aravinth Bheemaraj , Deepak Kumar , Dhiraj Daga , Saurabh Yadav , Srihari Nagaraj

This report evaluates the performance of text-in text-out Large Language Models (LLMs) to understand and generate Indic languages. This evaluation is used to identify and prioritize Indic languages suited for inclusion in safety benchmarks.…

计算与语言 · 计算机科学 2025-01-24 Aatman Vaidya , Tarunima Prabhakar , Denny George , Swair Shah

Automatic language identification is a natural language processing problem that tries to determine the natural language of a given content. In this paper we present a statistical method for automatic language identification of written text…

计算与语言 · 计算机科学 2018-06-15 Ciprian-Octavian Truică , Julien Velcin , Alexandru Boicea

While model architecture and training objectives are well-studied, tokenization, particularly in multilingual contexts, remains a relatively neglected aspect of Large Language Model (LLM) development. Existing tokenizers often exhibit high…

This work addresses the cross-corpora generalization issue for the low-resourced spoken language identification (LID) problem. We have conducted the experiments in the context of Indian LID and identified strikingly poor cross-corpora…

音频与语音处理 · 电气工程与系统科学 2023-03-02 Spandan Dey , Md Sahidullah , Goutam Saha

Automating the decision of whether a code change requires manual review is vital for maintaining software quality in modern development workflows. However, the emergence of new programming languages and frameworks creates a critical…

软件工程 · 计算机科学 2025-09-08 Yogev Cohen , Dudi Ohayon , Romy Somkin , Yehudit Aperstein , Alexander Apartsin

Spoken Language Identification (LID) is an important sub-task of Automatic Speech Recognition(ASR) that is used to classify the language(s) in an audio segment. Automatic LID plays an useful role in multilingual countries. In various…

音频与语音处理 · 电气工程与系统科学 2024-09-02 Parth Shastri , Chirag Patil , Poorval Wanere , Shrinivas Mahajan , Abhishek Bhatt , Hardik Sailor

Transfer learning is a popular strategy to improve the quality of low-resource machine translation. For an optimal transfer of the embedding layer, the child and parent model should share a substantial part of the vocabulary. This is not…

计算与语言 · 计算机科学 2020-10-01 Chantal Amrhein , Rico Sennrich

Large language models (LLMs) are increasingly applied in multilingual contexts, yet their capacity for consistent, logically grounded alignment across languages remains underexplored. We present a controlled evaluation framework for…

计算与语言 · 计算机科学 2025-08-21 Samir Abdaljalil , Erchin Serpedin , Khalid Qaraqe , Hasan Kurban

Language identification describes the task of recognizing the language of written text in documents. This information is crucial because it can be used to support the analysis of a document's vocabulary and context. Supervised learning…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Furkan Simsek , Brian Pfitzmann , Hendrik Raetz , Jona Otholt , Haojin Yang , Christoph Meinel

Context: Since it is well-established that developers spend a substantial portion of their time understanding source code, the ability to automatically identify algorithms within source code presents a valuable opportunity. This capability…

软件工程 · 计算机科学 2026-04-06 Denis Neumüller , Sebastian Boll , David Schüler , Matthias Tichy

Language identification from speech is a common preprocessing step in many spoken language processing systems. In recent years, this field has seen fast progress, mostly due to the use of self-supervised models pretrained on multilingual…

音频与语音处理 · 电气工程与系统科学 2022-07-04 Kunnar Kukk , Tanel Alumäe

This article describes an unsupervised language model adaptation approach that can be used to enhance the performance of language identification methods. The approach is applied to a current version of the HeLI language identification…

计算与语言 · 计算机科学 2019-03-27 Tommi Jauhiainen , Krister Lindén , Heidi Jauhiainen

Scene text recognition in low-resource Indian languages is challenging because of complexities like multiple scripts, fonts, text size, and orientations. In this work, we investigate the power of transfer learning for all the layers of deep…

计算机视觉与模式识别 · 计算机科学 2022-01-11 Sanjana Gunna , Rohit Saluja , C. V. Jawahar

Language-independent tokenisation (LIT) methods that do not require labelled language resources or lexicons have recently gained popularity because of their applicability in resource-poor languages. Moreover, they compactly represent a…

计算与语言 · 计算机科学 2020-02-26 Danushka Bollegala , Ryuichi Kiryo , Kosuke Tsujino , Haruki Yukawa

Language identification (LID) is an essential step in building high-quality multilingual datasets from web data. Existing LID tools (such as OpenLID or GlotLID) often struggle to identify closely related languages and to distinguish valid…

计算与语言 · 计算机科学 2026-02-24 Mariia Fedorova , Nikolay Arefyev , Maja Buljan , Jindřich Helcl , Stephan Oepen , Egil Rønningstad , Yves Scherrer

Language Identification (LID) systems are used to classify the spoken language from a given audio sample and are typically the first step for many spoken language processing tasks, such as Automatic Speech Recognition (ASR) systems. Without…

计算机视觉与模式识别 · 计算机科学 2017-08-17 Christian Bartz , Tom Herold , Haojin Yang , Christoph Meinel

The performance of Language Models (LMs) on low-resource, morphologically rich languages like Sinhala remains largely unexplored, particularly regarding script variation in digital communication. Sinhala exhibits script duality, with…

计算与语言 · 计算机科学 2026-05-11 Minuri Rajapakse , Ruvan Weerasinghe