中文
相关论文

相关论文: VoxPopuli: A Large-Scale Multilingual Speech Corpu…

200 篇论文

Large-scale generative models such as GPT and DALL-E have revolutionized the research community. These models not only generate high fidelity outputs, but are also generalists which can solve tasks not explicitly taught. In contrast, speech…

音频与语音处理 · 电气工程与系统科学 2023-10-20 Matthew Le , Apoorv Vyas , Bowen Shi , Brian Karrer , Leda Sari , Rashel Moritz , Mary Williamson , Vimal Manohar , Yossi Adi , Jay Mahadeokar , Wei-Ning Hsu

Speech foundation models have demonstrated exceptional capabilities in speech-related tasks. Nevertheless, these models often struggle with non-verbal audio data, such as vocalizations, baby crying, etc., which are critical for various…

音频与语音处理 · 电气工程与系统科学 2025-02-25 Alkis Koudounas , Moreno La Quatra , Marco Sabato Siniscalchi , Elena Baralis

The preservation of under-resourced languages requires digital tools and resources shaped by and for their speakers. We present the first dedicated ASR resources for Puno Quechua (ISO 639-3: qxp): (1) the largest speech corpus for any…

计算与语言 · 计算机科学 2026-05-28 Elwin Huaman , Adrian Gamarra Lafuente , Johanna Cordova , Anna Korhonen

Existing document-level machine translation resources are only available for a handful of languages, mostly high-resourced ones. To facilitate the training and evaluation of document-level translation and, more broadly, long-context…

计算与语言 · 计算机科学 2025-10-01 Dayyán O'Brien , Bhavitvya Malik , Ona de Gibert , Pinzhen Chen , Barry Haddow , Jörg Tiedemann

We present Multilingual Open Text (MOT), a new multilingual corpus containing text in 44 languages, many of which have limited existing text resources for natural language processing. The first release of the corpus contains over 2.8…

计算与语言 · 计算机科学 2022-06-10 Chester Palen-Michel , June Kim , Constantine Lignos

This paper introduces the Voices Obscured In Complex Environmental Settings (VOICES) corpus, a freely available dataset under Creative Commons BY 4.0. This dataset will promote speech and signal processing research of speech recorded by…

We present the Tongue and Lips corpus (TaL), a multi-speaker corpus of audio, ultrasound tongue imaging, and lip videos. TaL consists of two parts: TaL1 is a set of six recording sessions of one professional voice talent, a male native…

音频与语音处理 · 电气工程与系统科学 2020-11-20 Manuel Sam Ribeiro , Jennifer Sanger , Jing-Xuan Zhang , Aciel Eshky , Alan Wrench , Korin Richmond , Steve Renals

In order to simulate human language capacity, natural language processing systems must be able to reason about the dynamics of everyday situations, including their possible causes and effects. Moreover, they should be able to generalise the…

计算与语言 · 计算机科学 2020-10-28 Edoardo Maria Ponti , Goran Glavaš , Olga Majewska , Qianchu Liu , Ivan Vulić , Anna Korhonen

Contemporary works on abstractive text summarization have focused primarily on high-resource languages like English, mostly due to the limited availability of datasets for low/mid-resource ones. In this work, we present XL-Sum, a…

This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real online one-to-one…

计算与语言 · 计算机科学 2022-06-28 Chengfei Li , Shuhao Deng , Yaoping Wang , Guangjing Wang , Yaguang Gong , Changbin Chen , Jinfeng Bai

Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions of copyrighted or proprietary content, which raises…

Collecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data…

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound…

声音 · 计算机科学 2022-06-22 Yuan Gong , Jin Yu , James Glass

This paper introduces a new multi-speaker English dataset for training text-to-speech models. The dataset is based on LibriVox audiobooks and Project Gutenberg texts, both in the public domain. The new dataset contains about 292 hours of…

音频与语音处理 · 电气工程与系统科学 2021-06-16 Evelina Bakhturina , Vitaly Lavrukhin , Boris Ginsburg , Yang Zhang

Audio-driven talking head synthesis has achieved remarkable photorealism, yet state-of-the-art (SOTA) models exhibit a critical failure: they lack generalization to the full spectrum of human diversity in ethnicity, language, and age…

计算机视觉与模式识别 · 计算机科学 2025-08-20 Shunian Chen , Hejin Huang , Yexin Liu , Zihan Ye , Pengcheng Chen , Chenghao Zhu , Michael Guan , Rongsheng Wang , Junying Chen , Guanbin Li , Ser-Nam Lim , Harry Yang , Benyou Wang

We present a corpus of sentence-aligned triples of German audio, German text, and English translation, based on German audiobooks. The speech translation data consist of 110 hours of audio material aligned to over 50k parallel sentences. An…

计算与语言 · 计算机科学 2020-03-05 Benjamin Beilharz , Xin Sun , Sariya Karimova , Stefan Riezler

The advancement of vision-language models (VLMs) is hampered by a fragmented landscape of inconsistent and contaminated public datasets. We introduce FineVision, a meticulously collected, curated, and unified corpus of 24 million samples -…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Luis Wiedmann , Orr Zohar , Amir Mahla , Xiaohan Wang , Rui Li , Thibaud Frere , Leandro von Werra , Aritra Roy Gosthipaty , Andrés Marafioti

We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embeddings trained on these corpora. We create news article…

In this work, we introduce X-FACT: the largest publicly available multilingual dataset for factual verification of naturally existing real-world claims. The dataset contains short statements in 25 languages and is labeled for veracity by…

计算与语言 · 计算机科学 2021-06-18 Ashim Gupta , Vivek Srikumar