中文
相关论文

相关论文: Reconstructing Syllable Sequences in Abugida Scrip…

200 篇论文

When we transfer a pretrained language model to a new language, there are many axes of variation that change at once. To disentangle the impact of different factors like syntactic similarity and vocabulary similarity, we propose a set of…

计算与语言 · 计算机科学 2024-01-25 Zhengxuan Wu , Alex Tamkin , Isabel Papadimitriou

Despite being one of the most widely spoken languages globally, Bangla remains a low-resource language in the field of Natural Language Processing (NLP). Mainstream Automatic Speech Recognition (ASR) and Speaker Diarization systems for…

声音 · 计算机科学 2026-02-27 Zarif Ishmam , Zarif Mahir , Shafnan Wasif , Md. Ishtiak Moin

Handwritten character recognition is a crucial task because of its abundant applications. The recognition task of Bangla handwritten characters is especially challenging because of the cursive nature of Bangla characters and the presence of…

计算机视觉与模式识别 · 计算机科学 2024-02-06 Chandrika Saha , Md Mostafijur Rahman

Figurative language understanding remains a significant challenge for Large Language Models (LLMs), especially for low-resource languages. To address this, we introduce a new idiom dataset, a large-scale, culturally-grounded corpus of…

计算与语言 · 计算机科学 2026-02-16 Adib Sakhawat , Shamim Ara Parveen , Md Ruhul Amin , Shamim Al Mahmud , Md Saiful Islam , Tahera Khatun

Most Automatic Speech Recognition (ASR) systems formulate transcription as a prediction problem over orthographic units such as characters, subwords, or words. Although effective, such representations do not explicitly reflect the phonetic…

计算与语言 · 计算机科学 2026-05-28 Nghia Hieu Nguyen , Quan Ngoc Hoang , Long Hoang Huu Nguyen , Kiet Van Nguyen , Ngan Luu-Thuy Nguyen

In this work we investigate the ability of large language models to predict additive manufacturing defect regimes given a set of process parameter inputs. For this task we utilize a process parameter defect dataset to fine-tune a collection…

机器学习 · 计算机科学 2026-01-01 Peter Pak , Amir Barati Farimani

The choice of modeling units is crucial for automatic speech recognition (ASR) tasks. In mandarin scenarios, the Chinese characters represent meaning but are not directly related to the pronunciation. Thus only considering the writing of…

计算与语言 · 计算机科学 2022-10-19 Yuting Yang , Binbin Du , Yuke Li

Large Language Models (LLMs) based on transformer architectures have revolutionized a variety of domains, with tokenization playing a pivotal role in their pre-processing and fine-tuning stages. In multilingual models, particularly those…

计算与语言 · 计算机科学 2024-11-27 S. Tamang , D. J. Bora

We present Voxlect, a novel benchmark for modeling dialects and regional languages worldwide using speech foundation models. Specifically, we report comprehensive benchmark evaluations on dialects and regional language varieties in English,…

Financial fraud detection has emerged as a critical research challenge amid the rapid expansion of digital financial platforms. Although machine learning approaches have demonstrated strong performance in identifying fraudulent activities,…

Language models for agglutinative languages have always been hindered in past due to myriad of agglutinations possible to any given word through various affixes. We propose a method to diminish the problem of out-of-vocabulary words by…

计算与语言 · 计算机科学 2017-08-21 Seunghak Yu , Nilesh Kulkarni , Haejun Lee , Jihie Kim

Phonological reconstruction is one of the central problems in historical linguistics where a proto-word of an ancestral language is determined from the observed cognate words of daughter languages. Computational approaches to historical…

计算与语言 · 计算机科学 2023-12-27 V. S. D. S. Mahesh Akavarapu , Arnab Bhattacharya

This paper presents the work of restoring punctuation for ASR transcripts generated by multilingual ASR systems. The focus languages are English, Mandarin, and Malay which are three of the most popular languages in Singapore. To the best of…

计算与语言 · 计算机科学 2024-12-03 Abhinav Rao , Ho Thi-Nga , Chng Eng-Siong

Large Language Models (LLMs) have emerged as one of the most important breakthroughs in NLP for their impressive skills in language generation and other language-specific tasks. Though LLMs have been evaluated in various tasks, mostly in…

We investigate the problem of parsing conversational data of morphologically-rich languages such as Hindi where argument scrambling occurs frequently. We evaluate a state-of-the-art non-linear transition-based parsing system on a new…

计算与语言 · 计算机科学 2019-02-15 Riyaz Ahmad Bhat , Irshad Ahmad Bhat , Dipti Misra Sharma

Aspect-based sentiment analysis (ABSA) has made significant strides, yet challenges remain for low-resource languages due to the predominant focus on English. Current cross-lingual ABSA studies often centre on simpler tasks and rely heavily…

计算与语言 · 计算机科学 2025-08-15 Jakub Šmíd , Pavel Přibáň , Pavel Král

Word segmentation is a fundamental pre-processing step for Thai Natural Language Processing. The current off-the-shelf solutions are not benchmarked consistently, so it is difficult to compare their trade-offs. We conducted a speed and…

计算与语言 · 计算机科学 2019-11-19 Pattarawat Chormai , Ponrawee Prasertsom , Attapol Rutherford

This paper presents a constraint-based morphological disambiguation approach that is applicable languages with complex morphology--specifically agglutinative languages with productive inflectional and derivational morphological phenomena.…

cmp-lg · 计算机科学 2008-02-03 Kemal Oflazer , Gokhan Tur

Bengali remains a low-resource language in speech technology, especially for complex tasks like long-form transcription and speaker diarization. This paper presents a multistage approach developed for the "DL Sprint 4.0 - Bengali Long-Form…

声音 · 计算机科学 2026-03-04 Epshita Jahan , Khandoker Md Tanjinul Islam , Pritom Biswas , Tafsir Al Nafin

Bangla is the 7th most widely spoken language globally, with a staggering 234 million native speakers primarily hailing from India and Bangladesh. This morphologically rich language boasts a rich literary tradition, encompassing diverse…

计算与语言 · 计算机科学 2023-10-19 Saumajit Saha , Albert Nanda