中文
相关论文

相关论文: Uzbek Cyrillic-Latin-Cyrillic Machine Transliterat…

200 篇论文

Document-level machine translation focuses on the translation of entire documents from a source to a target language. It is widely regarded as a challenging task since the translation of the individual sentences in the document needs to…

计算与语言 · 计算机科学 2020-10-21 Inigo Jauregi Unanue , Nazanin Esmaili , Gholamreza Haffari , Massimo Piccardi

Large language models fragment Kazakh text into many more tokens than equivalent English text, because their tokenizers were built for high-resource languages. This tokenizer tax inflates compute, shortens the effective context window, and…

计算与语言 · 计算机科学 2026-03-31 Rauan Akylzhanov

This paper describes the submission of the UPC Machine Translation group to the IWSLT 2023 Offline Speech Translation task. Our Speech Translation systems utilize foundation models for speech (wav2vec 2.0) and text (mBART50). We incorporate…

计算与语言 · 计算机科学 2023-06-05 Ioannis Tsiamas , Gerard I. Gállego , José A. R. Fonollosa , Marta R. Costa-jussà

Most of the recent work on terminology integration in machine translation has assumed that terminology translations are given already inflected in forms that are suitable for the target language sentence. In day-to-day work of professional…

计算与语言 · 计算机科学 2021-01-26 Toms Bergmanis , Mārcis Pinnis

This article discusses the problem of handwriting recognition in Kazakh and Russian languages. This area is poorly studied since in the literature there are almost no works in this direction. We have tried to describe various approaches and…

计算机视觉与模式识别 · 计算机科学 2021-02-10 Daniyar Nurseitov , Kairat Bostanbekov , Maksat Kanatov , Anel Alimova , Abdelrahman Abdallah , Galymzhan Abdimanap

Despite impressive empirical successes of neural machine translation (NMT) on standard benchmarks, limited parallel data impedes the application of NMT models to many language pairs. Data augmentation methods such as back-translation make…

计算与语言 · 计算机科学 2019-10-08 Chunting Zhou , Xuezhe Ma , Junjie Hu , Graham Neubig

Machine translation for low resource language pairs is a challenging task. This task could become extremely difficult once a speaker uses code switching. We propose a method to build a machine translation model for code-switched…

计算与语言 · 计算机科学 2025-03-27 Maksim Borisov , Zhanibek Kozhirbayev , Valentin Malykh

This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the performance of large foundation models against a multimodal…

计算与语言 · 计算机科学 2026-02-09 Yu Wu , Ke Shu , Jonas Fischer , Lidia Pivovarova , David Rosson , Eetu Mäkelä , Mikko Tolonen

In this paper, we focus on the challenge of learning controllable text simplifications in unsupervised settings. While this problem has been previously discussed for supervised learning algorithms, the literature on the analogies in…

计算与语言 · 计算机科学 2020-12-04 Oleg Kariuk , Dima Karamshuk

Eye tracking data during reading is a useful source of information to understand the cognitive processes that take place during language comprehension processes. Different languages account for different brain triggers , however there seems…

计算与语言 · 计算机科学 2022-03-31 Harshvardhan Srivastava

To make robots accessible to a broad audience, it is critical to endow them with the ability to take universal modes of communication, like commands given in natural language, and extract a concrete desired task specification, defined using…

计算与语言 · 计算机科学 2023-03-22 Jiayi Pan , Glen Chou , Dmitry Berenson

Data contamination undermines the validity of Large Language Model evaluation by enabling models to rely on memorized benchmark content rather than true generalization. While prior work has proposed contamination detection methods, these…

计算与语言 · 计算机科学 2026-01-22 Chaymaa Abbas , Nour Shamaa , Mariette Awad

This paper provides an analysis of character-level machine translation models used in pivot-based translation when applied to sparse and noisy datasets, such as crowdsourced movie subtitles. In our experiments, we find that such…

计算与语言 · 计算机科学 2021-09-29 Jörg Tiedemann , Preslav Nakov

Most large language models are fine-tuned using either expensive human-annotated data or GPT-4 generated data which cannot guarantee performance in certain domains. We argue that although the web-crawled data often has formatting errors…

计算与语言 · 计算机科学 2024-08-16 Jing Zhou , Chenglin Jiang , Wei Shen , Xiao Zhou , Xiaonan He

Chinese Spell Checking (CSC) aims to detect and correct spelling errors in sentences. Despite Large Language Models (LLMs) exhibit robust capabilities and are widely applied in various tasks, their performance on CSC is often…

计算与语言 · 计算机科学 2024-10-29 Kunting Li , Yong Hu , Liang He , Fandong Meng , Jie Zhou

Language identification is an important Natural Language Processing task. It has been thoroughly researched in the literature. However, some issues are still open. This work addresses the identification of the related low-resource languages…

计算与语言 · 计算机科学 2022-03-10 Olha Dovbnia , Anna Wróblewska

Large language models (LLMs) are increasingly used to assist scientists across diverse workflows. A key challenge is generating high-quality figures from textual descriptions, often represented as TikZ programs that can be rendered as…

人工智能 · 计算机科学 2026-03-26 Christian Greisinger , Steffen Eger

Handwritten text recognition and optical character recognition solutions show excellent results with processing data of modern era, but efficiency drops with Latin documents of medieval times. This paper presents a deep learning method to…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Maksym Voloshchuk , Bohdana Zarembovska , Mykola Kozlenko

This paper describes the solutions submitted by the UPB team to the AuTexTification shared task, featured as part of IberLEF-2023. Our team participated in the first subtask, identifying text documents produced by large language models…

计算与语言 · 计算机科学 2023-08-04 Andrei-Alexandru Preda , Dumitru-Clementin Cercel , Traian Rebedea , Costin-Gabriel Chiru

As research on machine translation moves to translating text beyond the sentence level, it remains unclear how effective automatic evaluation metrics are at scoring longer translations. In this work, we first propose a method for creating…

计算与语言 · 计算机科学 2023-08-29 Daniel Deutsch , Juraj Juraska , Mara Finkelstein , Markus Freitag