中文
相关论文

相关论文: MOROCO: The Moldavian and Romanian Dialectal Corpu…

200 篇论文

This paper proposes a method for extracting a lightweight subset from a text-to-speech (TTS) corpus ensuring synthetic speech quality. In recent years, methods have been proposed for constructing large-scale TTS corpora by collecting…

声音 · 计算机科学 2023-09-18 Kentaro Seki , Shinnosuke Takamichi , Takaaki Saeki , Hiroshi Saruwatari

The CMU Wilderness Multilingual Speech Dataset (Black, 2019) is a newly published multilingual speech dataset based on recorded readings of the New Testament. It provides data to build Automatic Speech Recognition (ASR) and Text-to-Speech…

计算与语言 · 计算机科学 2020-02-27 Marcely Zanon Boito , William N. Havard , Mahault Garnerin , Éric Le Ferrand , Laurent Besacier

Parliamentary transcripts provide a valuable resource to understand the reality and know about the most important facts that occur over time in our societies. Furthermore, the political debates captured in these transcripts facilitate…

Large multimodal models trained on natural documents, which interleave images and text, outperform models trained on image-text pairs on various multimodal benchmarks. However, the datasets used to train these models have not been released,…

Multilingual topic models enable crosslingual tasks by extracting consistent topics from multilingual corpora. Most models require parallel or comparable training corpora, which limits their ability to generalize. In this paper, we first…

计算与语言 · 计算机科学 2018-06-13 Shudong Hao , Michael J. Paul

Lombard, an underresourced language variety spoken by approximately 3.8 million people in Northern Italy and Southern Switzerland, lacks a unified orthographic standard. Multiple orthographic systems exist, creating challenges for NLP…

计算与语言 · 计算机科学 2026-03-31 Edoardo Signoroni , Pavel Rychlý

We address a notable gap in Natural Language Processing (NLP) by introducing a collection of resources designed to improve Machine Translation (MT) for low-resource languages, with a specific focus on African languages. First, we introduce…

计算与语言 · 计算机科学 2024-07-15 AbdelRahim Elmadany , Ife Adebara , Muhammad Abdul-Mageed

Topic modelling is a pivotal unsupervised machine learning technique for extracting valuable insights from large document collections. Existing neural topic modelling methods often encode contextual information of documents, while ignoring…

计算与语言 · 计算机科学 2025-02-07 Yanan Ma , Chenghao Xiao , Chenhan Yuan , Sabine N van der Veer , Lamiece Hassan , Chenghua Lin , Goran Nenadic

We introduce a new dataset of conversational speech representing English from India, Nigeria, and the United States. The Multi-Dialect Dataset of Dialogues (MD3) strikes a new balance between open-ended conversational speech and…

计算与语言 · 计算机科学 2023-05-22 Jacob Eisenstein , Vinodkumar Prabhakaran , Clara Rivera , Dorottya Demszky , Devyani Sharma

We introduce MULTI-EURLEX, a new multilingual dataset for topic classification of legal documents. The dataset comprises 65k European Union (EU) laws, officially translated in 23 languages, annotated with multiple labels from the EUROVOC…

计算与语言 · 计算机科学 2021-09-08 Ilias Chalkidis , Manos Fergadiotis , Ion Androutsopoulos

Speech neuroprosthesis systems decode intended speech from neural activity in the absence of audible output, offering a path to restoring communication for individuals with speech-impairing conditions. Current approaches decode…

神经元与认知 · 定量生物学 2026-05-28 Yuanhao Chen , Peter Chin

We present a preview of the Syntactic Acceptability Dataset, a resource being designed for both syntax and computational linguistics research. In its current form, the dataset comprises 1,000 English sequences from the syntactic discourse:…

计算与语言 · 计算机科学 2025-06-24 Tom S Juzek

We introduce Morse, a recurrent encoder-decoder model that produces morphological analyses of each word in a sentence. The encoder turns the relevant information about the word and its context into a fixed size vector representation and the…

计算与语言 · 计算机科学 2019-09-25 Ekin Akyürek , Erenay Dayanık , Deniz Yuret

We propose a generative framework for multi-track music source separation (MSS) that reformulates the task as conditional discrete token generation. Unlike conventional approaches that directly estimate continuous signals in the time or…

音频与语音处理 · 电气工程与系统科学 2026-04-20 Pengbo Lyu , Xiangyu Zhao , Chengwei Liu , Haoyin Yan , Xiaotao Liang , Hongyu Wang , Shaofei Xue

The rapid advancement of Large Language Models (LLMs) has transformed various domains, particularly computer science (CS) education. These models exhibit remarkable capabilities in code-related tasks and problem-solving, raising questions…

计算机与社会 · 计算机科学 2025-09-30 Dumitran Adrian Marius , Theodor-Pierre Moroianu , Buca Mihnea-Vicentiu

This article examines semantic shifts in psychological concepts across scientific and popular media discourse using methods of distributional semantics applied to Russian-language corpora. Two corpora were compiled: a scientific corpus of…

计算与语言 · 计算机科学 2026-04-02 Orlova Anastasia

Lip reading or visual speech recognition has gained significant attention in recent years, particularly because of hardware development and innovations in computer vision. While considerable progress has been obtained, most models have only…

计算机视觉与模式识别 · 计算机科学 2023-10-10 Emilian-Claudiu Mănescu , Răzvan-Alexandru Smădu , Andrei-Marius Avram , Dumitru-Clementin Cercel , Florin Pop

SiPaKosa is a comprehensive corpus of Sinhala and Pali doctrinal texts comprising approximately 786K sentences and 9.25M words, incorporating 16 copyright-cleared historical Buddhist documents alongside the complete web-scraped Tripitaka…

计算与语言 · 计算机科学 2026-04-01 Ranidu Gurusinghe , Nevidu Jayatilleke

We present "Testimole-conversational" a massive collection of discussion boards messages in the Italian language. The large size of the corpus, more than 30B word-tokens (1996-2024), renders it an ideal dataset for native Italian Large…

计算与语言 · 计算机科学 2026-04-10 Matteo Rinaldi , Rossella Varvara , Viviana Patti