中文
相关论文

相关论文: SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.…

200 篇论文

In this paper we describe a novel framework for the discovery of the topical content of a data corpus, and the tracking of its complex structural changes across the temporal dimension. In contrast to previous work our model does not impose…

信息检索 · 计算机科学 2015-02-10 Adham Beykikhoshk , Ognjen Arandjelovic , Dinh Phung , Svetha Venkatesh

Visually Rich Documents (VRDs), comprising elements such as charts, tables, and paragraphs, convey complex information across diverse domains. However, extracting key information from these documents remains labour-intensive, particularly…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Yihao Ding , Soyeon Caren Han , Zechuan Li , Hyunsuk Chung

Sanskrit, an ancient language with a rich linguistic heritage, presents unique challenges for automatic speech recognition (ASR) due to its phonemic complexity and the phonetic transformations that occur at word junctures, similar to the…

计算与语言 · 计算机科学 2025-06-03 Sujeet Kumar , Pretam Ray , Abhinay Beerukuri , Shrey Kamoji , Manoj Balaji Jagadeeshan , Pawan Goyal

Extracting concise information from scientific documents aids learners, researchers, and practitioners. Automatic Text Summarization (ATS), a key Natural Language Processing (NLP) application, automates this process. While ATS methods exist…

计算与语言 · 计算机科学 2025-04-22 Rondik Hadi Abdulrahman , Hossein Hassani

We introduce ChronoQA, a large-scale benchmark dataset for Chinese question answering, specifically designed to evaluate temporal reasoning in Retrieval-Augmented Generation (RAG) systems. ChronoQA is constructed from over 300,000 news…

计算与语言 · 计算机科学 2025-08-19 Ziyang Chen , Erxue Min , Xiang Zhao , Yunxin Li , Xin Jia , Jinzhi Liao , Jichao Li , Shuaiqiang Wang , Baotian Hu , Dawei Yin

We present 20min-XD (20 Minuten cross-lingual document-level), a French-German, document-level comparable corpus of news articles, sourced from the Swiss online news outlet 20 Minuten/20 minutes. Our dataset comprises around 15,000 article…

计算与语言 · 计算机科学 2025-05-01 Michelle Wastl , Jannis Vamvas , Selena Calleri , Rico Sennrich

We introduce a data set called DCH-2, which contains 4,390 real customer-helpdesk dialogues in Chinese and their English translations. DCH-2 also contains dialogue-level annotations and turn-level annotations obtained independently from…

计算与语言 · 计算机科学 2021-06-01 Zhaohao Zeng , Tetsuya Sakai

This paper presents a large-scale corpus for non-task-oriented dialogue response selection, which contains over 27K distinct prompts more than 82K responses collected from social media. To annotate this corpus, we define a 5-grade rating…

计算与语言 · 计算机科学 2018-05-16 Jing Li , Yan Song , Haisong Zhang , Shuming Shi

This paper reviews the state-of-the-art of semantic change computation, one emerging research field in computational linguistics, proposing a framework that summarizes the literature by identifying and expounding five essential components…

计算与语言 · 计算机科学 2018-06-20 Xuri Tang

We present RUSLAN -- a new open Russian spoken language corpus for the text-to-speech task. RUSLAN contains 22200 audio samples with text annotations -- more than 31 hours of high-quality speech of one person -- being the largest annotated…

音频与语音处理 · 电气工程与系统科学 2019-06-28 Lenar Gabdrakhmanov , Rustem Garaev , Evgenii Razinkov

Large language models (LLMs) use data to learn about the world in order to produce meaningful correlations and predictions. As such, the nature, scale, quality, and diversity of the datasets used to train these models, or to support their…

We present a comprehensive corpus of Russian primary and secondary legislation adopted between 1991 and 2025, comprising 304,382 texts (194,425,905 tokens). The corpus is available in two versions: the basic version contains texts with…

计算与语言 · 计算机科学 2026-04-29 Denis Saveliev , Ruslan Kuchakov

The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap, we introduce…

This article investigates how translation memories (TM) can be created by translators or other language professionals in order to compile domain-specific parallel corpora , which can then be used in different scenarios, such as machine…

计算与语言 · 计算机科学 2024-09-05 Gokhan Dogru

Sanskrit Subhasitas encapsulate centuries of cultural and philosophical wisdom, yet remain underutilized in the digital age due to linguistic and contextual barriers. In this work, we present Pragya, a retrieval-augmented generation (RAG)…

计算与语言 · 计算机科学 2026-01-13 Tanisha Raorane , Prasenjit Kole

Speech translation for Indian languages remains a challenging task due to the scarcity of large-scale, publicly available datasets that capture the linguistic diversity and domain coverage essential for real-world applications. Existing…

This paper presents the first publicly available version of the Carolina Corpus and discusses its future directions. Carolina is a large open corpus of Brazilian Portuguese texts under construction using web-as-corpus methodology enhanced…

As natural language corpora expand at an unprecedented rate, manual annotation remains a significant methodological bottleneck in corpus linguistic work. We address this challenge by presenting a scalable pipeline for automating grammatical…

计算与语言 · 计算机科学 2026-02-11 Cameron Morin , Matti Marttinen Larsson

In the process of numerically modeling natural languages, developing language embeddings is a vital step. However, it is challenging to develop functional embeddings for resource-poor languages such as Sinhala, for which sufficiently large…

计算与语言 · 计算机科学 2022-10-27 Gihan Weeraprameshwara , Vihanga Jayawickrama , Nisansa de Silva , Yudhanjaya Wijeratne
‹ 上一页 1 8 9 10 下一页 ›