中文
相关论文

相关论文: Uzbek Cyrillic-Latin-Cyrillic Machine Transliterat…

200 篇论文

This study tackles unsupervised domain adaptation of reading comprehension (UDARC). Reading comprehension (RC) is a task to learn the capability for question answering with textual sources. State-of-the-art models on RC still do not have…

计算与语言 · 计算机科学 2020-05-22 Kosuke Nishida , Kyosuke Nishida , Itsumi Saito , Hisako Asano , Junji Tomita

This article presents an original method for Text-to-Sign Translation. It compensates data scarcity using a domain-specific parallel corpus of alignments between text and hierarchical formal descriptions of Sign Language videos in AZee.…

计算与语言 · 计算机科学 2022-05-09 Élise Bertin-Lemée , Annelies Braffort , Camille Challant , Claire Danet , Michael Filhol

Informal transliteration from other languages to English is prevalent in social media threads, instant messaging, and discussion forums. Without identifying the language of such transliterated text, users who do not speak that language…

音频与语音处理 · 电气工程与系统科学 2020-04-29 Sourav Sen

Cross-lingual text classification(CLTC) is the task of classifying documents written in different languages into the same taxonomy of categories. This paper presents a novel approach to CLTC that builds on model distillation, which adapts…

计算与语言 · 计算机科学 2018-03-29 Ruochen Xu , Yiming Yang

Natural language understanding systems struggle with low-resource languages, including many dialects of high-resource ones. Dialect-to-standard normalization attempts to tackle this issue by transforming dialectal text so that it can be…

计算与语言 · 计算机科学 2025-06-11 Antonios Dimakis , John Pavlopoulos , Antonios Anastasopoulos

Machine translation has gained much attention in recent years. It is a sub-field of computational linguistic which focus on translating text from one language to other language. Among different translation techniques, neural network…

计算与语言 · 计算机科学 2020-01-16 Usman Mohy ud Din

Today text classification becomes critical task for concerned individuals for numerous purposes. Hence, several researches have been conducted to develop automatic text classification for national and international languages. However, the…

计算与语言 · 计算机科学 2023-05-09 Jawid Ahmad Baktash , Mursal Dawodi , Mohammad Zarif Joya , Nematullah Hassanzada

Machine translation (MT) of user-generated content (UGC) poses unique challenges, including handling slang, emotion, and literary devices like irony and sarcasm. Evaluating the quality of these translations is challenging as current metrics…

计算与语言 · 计算机科学 2024-10-07 Shenbin Qian , Constantin Orăsan , Diptesh Kanojia , Félix do Carmo

Text classification systems will help to solve the text clustering problem in the Azerbaijani language. There are some text-classification applications for foreign languages, but we tried to build a newly developed system to solve this…

计算与语言 · 计算机科学 2020-01-01 Umid Suleymanov , Behnam Kiani Kalejahi , Elkhan Amrahov , Rashid Badirkhanli

Kazakh, a Turkic language spoken by over 22 million people, remains underserved by existing multilingual language models, which allocate minimal capacity to low-resource languages and employ tokenizers ill-suited to agglutinative…

计算与语言 · 计算机科学 2026-03-24 Saken Tukenov

The popularity of automatic speech-to-speech translation for human conversations is growing, but the quality varies significantly depending on the language pair. In a context of community interpreting for low-resource languages, namely…

计算与语言 · 计算机科学 2025-06-03 Andrei Popescu-Belis , Alexis Allemann , Teo Ferrari , Gopal Krishnamani

Understanding the qualitative intent of citations is essential for a comprehensive assessment of academic research, a task that poses unique challenges for agglutinative languages like Turkish. This paper introduces a systematic methodology…

计算与语言 · 计算机科学 2025-11-04 Kemal Sami Karaca , Bahaeddin Eravcı

Tibetan is a low-resource language. In order to alleviate the shortage of parallel corpus between Tibetan and Chinese, this paper uses two monolingual corpora and a small number of seed dictionaries to learn the semi-supervised method with…

计算与语言 · 计算机科学 2021-10-05 Enshuai Hou , Jie zhu

Neural machine translation has become a major alternative to widely used phrase-based statistical machine translation. We notice however that much of research on neural machine translation has focused on European languages despite its…

计算与语言 · 计算机科学 2016-06-09 Amjad Almahairi , Kyunghyun Cho , Nizar Habash , Aaron Courville

Evaluating machine translation (MT) for low-resource languages poses a persistent challenge, primarily due to the limited availability of high quality reference translations. This issue is further exacerbated in languages with multiple…

计算与语言 · 计算机科学 2025-05-20 Md. Atiqur Rahman , Sabrina Islam , Mushfiqul Haque Omi

Diacritization plays a pivotal role in improving readability and disambiguating the meaning of Arabic texts. Efforts have so far focused on marking every eligible character (Full Diacritization). Comparatively overlooked, Partial…

计算与语言 · 计算机科学 2024-08-12 Muhammad ElNokrashy , Badr AlKhamissi

Although multilingual LLMs have achieved remarkable performance across benchmarks, we find they continue to underperform on non-Latin script languages across contemporary LLM families. This discrepancy arises from the fact that LLMs are…

计算与语言 · 计算机科学 2025-06-27 Hoang H Nguyen , Khyati Mahajan , Vikas Yadav , Julian Salazar , Philip S. Yu , Masoud Hashemi , Rishabh Maheshwary

We introduce a simple approach that uses a large language model (LLM) to automatically implement a fully interpretable rule-based data-to-text system in pure Python. Experimental evaluation on the WebNLG dataset showed that such a…

计算与语言 · 计算机科学 2025-03-03 Jędrzej Warczyński , Mateusz Lango , Ondrej Dusek

This paper presents the first comprehensive study on automatic readability assessment of Turkish texts. We combine state-of-the-art neural network models with linguistic features at lexical, morphological, syntactic and discourse levels to…

计算与语言 · 计算机科学 2025-09-05 Ahmet Yavuz Uluslu , Gerold Schneider

Modern language models are internally -- and mathematically -- distributions over $\it{token}$ strings rather than $\it{character}$ strings, posing numerous challenges for programmers building user applications on top of them. For example,…