中文
相关论文

相关论文: Jambu: A historical linguistic database for South …

200 篇论文

This paper introduces AFRIDOC-MT, a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yor\`ub\'a, and Zulu. The dataset comprises 334 health and 271 information…

The VT legacy system, comprising approximately 2.5 million lines of PL/SQL code, lacks consistent documentation and automated tests, posing significant challenges for refactoring and modernisation. This study investigates the feasibility of…

软件工程 · 计算机科学 2025-08-28 Lola Solovyeva , Eduardo Carneiro Oliveira , Shiyu Fan , Alper Tuncay , Shamil Gareev , Andrea Capiluppi

Hindi, one of the most spoken language of India, exhibits a diverse array of accents due to its usage among individuals from diverse linguistic origins. To enable a robust evaluation of Hindi ASR systems on multiple accents, we create a…

计算与语言 · 计算机科学 2024-08-22 Tahir Javed , Janki Nawale , Sakshi Joshi , Eldho George , Kaushal Bhogale , Deovrat Mehendale , Mitesh M. Khapra

The Bangla language is the seventh most spoken language, with 265 million native and non-native speakers worldwide. However, English is the predominant language for online resources and technical knowledge, journals, and documentation.…

Text embeddings are an essential building component of several NLP tasks such as retrieval-augmented generation which is crucial for preventing hallucinations in LLMs. Despite the recent release of massively multilingual MTEB (MMTEB),…

计算与语言 · 计算机科学 2026-03-09 Kosei Uemura , Miaoran Zhang , David Ifeoluwa Adelani

In this paper, we introduce the MLM (Multiple Languages and Modalities) dataset - a new resource to train and evaluate multitask systems on samples in multiple modalities and three languages. The generation process and inclusion of semantic…

机器学习 · 计算机科学 2020-10-27 Jason Armitage , Endri Kacupaj , Golsa Tahmasebzadeh , Swati , Maria Maleshkova , Ralph Ewerth , Jens Lehmann

While Large Language Models (LLMs) have significantly advanced Text-to-SQL performance, existing benchmarks predominantly focus on Western contexts and simplified schemas, leaving a gap in real-world, non-Western applications. We present…

计算与语言 · 计算机科学 2026-04-16 Aviral Dawar , Roshan Karanth , Vikram Goyal , Dhruv Kumar

We introduce INDOTABVQA, a benchmark for evaluating cross-lingual Table Visual Question Answering (VQA) on real-world document images in Bahasa Indonesia. The dataset comprises 1,593 document images across three visual styles (bordered,…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Somraj Gautam , Anathapindika Dravichi , Gaurav Harit

Designing reliable Speech Emotion Recognition systems is a complex task that inevitably requires sufficient data for training purposes. Such extensive datasets are currently available in only a few languages, including English, German, and…

Access to humanities research databases is often hindered by the limitations of traditional interaction formats, particularly in the methods of searching and response generation. This study introduces an LLM-based smart assistant designed…

计算与语言 · 计算机科学 2025-06-03 Alexander Sergeev , Valeriya Goloviznina , Mikhail Melnichenko , Evgeny Kotelnikov

Contemporary database systems, while effective, suffer severe issues related to complexity and usability, especially among individuals who lack technical expertise but are unfamiliar with query languages like Structured Query Language…

数据库 · 计算机科学 2025-07-25 M. Tedeschi , S. Rizwan , C. Shringi , V. Devram Chandgir , S. Belich

We present TMMLU+, a new benchmark designed for Traditional Chinese language understanding. TMMLU+ is a multi-choice question-answering dataset with 66 subjects from elementary to professional level. It is six times larger and boasts a more…

计算与语言 · 计算机科学 2024-07-12 Zhi-Rui Tam , Ya-Ting Pai , Yen-Wei Lee , Jun-Da Chen , Wei-Min Chu , Sega Cheng , Hong-Han Shuai

The recent advances in deep-learning have led to the development of highly sophisticated systems with an unquenchable appetite for data. On the other hand, building good deep-learning models for low-resource languages remains a challenging…

计算与语言 · 计算机科学 2024-02-20 Maithili Sabane , Onkar Litake , Aman Chadha

We conducted a labeling work on a spoken Japanese dataset (I-JAS) for the text classification, which contains 50 interview dialogues of two-way Japanese conversation that discuss the participants' past present and future. Each dialogue is…

计算与语言 · 计算机科学 2021-03-23 Changzeng Fu

Large Language Models (LLMs) have tremendous potential to play a key role in supporting mathematical reasoning, with growing use in education and AI research. However, most existing benchmarks are limited to English, creating a significant…

计算机与社会 · 计算机科学 2025-10-16 Tabia Tanzin Prama , Christopher M. Danforth , Peter Sheridan Dodds

Speech translation for Indian languages remains a challenging task due to the scarcity of large-scale, publicly available datasets that capture the linguistic diversity and domain coverage essential for real-world applications. Existing…

This paper describes the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages. The dataset includes, for each language: 1) native script Wikipedia text; 2) a romanization…

计算与语言 · 计算机科学 2020-07-03 Brian Roark , Lawrence Wolf-Sonkin , Christo Kirov , Sabrina J. Mielke , Cibu Johny , Isin Demirsahin , Keith Hall

This paper presents the submission by the CMU-01 team to the SIGMORPHON 2019 task 2 of Morphological Analysis and Lemmatization in Context. This task requires us to produce the lemma and morpho-syntactic description of each token in a…

计算与语言 · 计算机科学 2019-07-25 Aditi Chaudhary , Elizabeth Salesky , Gayatri Bhat , David R. Mortensen , Jaime G. Carbonell , Yulia Tsvetkov

Recently, Large Language Models (LLMs) have dominated much of the artificial intelligence scene with their ability to process and generate natural languages. However, the majority of LLM research and development remains English-centric,…