中文
相关论文

相关论文: The ParlaSpeech Collection of Automatically Genera…

200 篇论文

Parliamentary speech generation presents specific challenges for large language models beyond standard text generation tasks. Unlike general text generation, parliamentary speeches require not only linguistic quality but also political…

计算与语言 · 计算机科学 2025-11-12 Marios Koniaris , Argyro Tsipi , Panayiotis Tsanakas

This paper describes an English audio and textual dataset of debating speeches, a unique resource for the growing research field of computational argumentation and debating technologies. We detail the process of speech recording by…

Existing conversational datasets consist either of written proxies for dialog or small-scale transcriptions of natural speech. We introduce 'Interview': a large-scale (105K conversations) media dialog dataset collected from news interview…

计算与语言 · 计算机科学 2020-04-08 Bodhisattwa Prasad Majumder , Shuyang Li , Jianmo Ni , Julian McAuley

Domain-specific data is the crux of the successful transfer of machine learning systems from benchmarks to real life. In simple problems such as image classification, crowdsourcing has become one of the standard tools for cheap and…

声音 · 计算机科学 2021-10-22 Nikita Pavlichenko , Ivan Stelmakh , Dmitry Ustalov

We present ASR Bundestag, a dataset for automatic speech recognition in German, consisting of 610 hours of aligned audio-transcript pairs for supervised training as well as 1,038 hours of unlabeled audio snippets for self-supervised…

计算与语言 · 计算机科学 2023-02-14 Johannes Wirth , René Peinl

Political discourse datasets are important for gaining political insights, analyzing communication strategies or social science phenomena. Although numerous political discourse corpora exist, comprehensive, high-quality, annotated datasets…

The success of large language models has driven interest in developing similar speech processing capabilities. However, a key challenge is the scarcity of high-quality spontaneous speech data, as most existing datasets contain scripted…

Multi-task and multilingual approaches benefit large models, yet speech processing for low-resource languages remains underexplored due to data scarcity. To address this, we present Granary, a large-scale collection of speech datasets for…

This paper introduces GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and…

Speech datasets available in the public domain are often underutilized because of challenges in discoverability and interoperability. A comprehensive framework has been designed to survey, catalog, and curate available speech datasets,…

音频与语音处理 · 电气工程与系统科学 2024-08-02 Michał Junczyk

Public sources like parliament meeting recordings and transcripts provide ever-growing material for the training and evaluation of automatic speech recognition (ASR) systems. In this paper, we publish and analyse the Finnish parliament ASR…

计算与语言 · 计算机科学 2022-03-29 Anja Virkkunen , Aku Rouhe , Nhan Phan , Mikko Kurimo

Recent studies in speech-driven 3D talking head generation have achieved convincing results in verbal articulations. However, generating accurate lip-syncs degrades when applied to input speech in other languages, possibly due to the lack…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Kim Sung-Bin , Lee Chae-Yeon , Gihun Son , Oh Hyun-Bin , Janghoon Ju , Suekyeong Nam , Tae-Hyun Oh

This paper presents SaSLaW, a spontaneous dialogue speech corpus containing synchronous recordings of what speakers speak, listen to, and watch. Humans consider the diverse environmental factors and then control the features of their…

音频与语音处理 · 电气工程与系统科学 2024-08-14 Osamu Take , Shinnosuke Takamichi , Kentaro Seki , Yoshiaki Bando , Hiroshi Saruwatari

This paper introduces Multilingual LibriSpeech (MLS) dataset, a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages, including about 44.5K hours of…

音频与语音处理 · 电气工程与系统科学 2020-12-22 Vineel Pratap , Qiantong Xu , Anuroop Sriram , Gabriel Synnaeve , Ronan Collobert

Using supervised automatic summarisation methods requires sufficient corpora that include pairs of documents and their summaries. Similarly to many tasks in natural language processing, most of the datasets available for summarization are…

Recent works in spoken language translation (SLT) have attempted to build end-to-end speech-to-text translation without using source language transcription during learning or decoding. However, while large quantities of parallel texts (such…

计算与语言 · 计算机科学 2018-02-12 Ali Can Kocabiyikoglu , Laurent Besacier , Olivier Kraif

Parliamentary transcripts provide a valuable resource to understand the reality and know about the most important facts that occur over time in our societies. Furthermore, the political debates captured in these transcripts facilitate…

The Norwegian Parliamentary Speech Corpus (NPSC) is a speech dataset with recordings of meetings from Stortinget, the Norwegian parliament. It is the first, publicly available dataset containing unscripted, Norwegian speech designed for…

计算与语言 · 计算机科学 2023-02-08 Per Erik Solberg , Pablo Ortiz

State-of-the-art performance for Automatic Speech Recognition (ASR) largely depends on the availability of large-scale labeled corpora. This creates a demand for increased data collection efforts, particularly for under-represented…

The CMU Wilderness Multilingual Speech Dataset (Black, 2019) is a newly published multilingual speech dataset based on recorded readings of the New Testament. It provides data to build Automatic Speech Recognition (ASR) and Text-to-Speech…

计算与语言 · 计算机科学 2020-02-27 Marcely Zanon Boito , William N. Havard , Mahault Garnerin , Éric Le Ferrand , Laurent Besacier