中文
相关论文

相关论文: 10 hours data is all you need

200 篇论文

Previous fake speech datasets were constructed from a defender's perspective to develop countermeasure (CM) systems without considering diverse motivations of attackers. To better align with real-life scenarios, we created…

音频与语音处理 · 电气工程与系统科学 2025-01-07 Hieu-Thi Luong , Haoyang Li , Lin Zhang , Kong Aik Lee , Eng Siong Chng

We present improvements in automatic speech recognition (ASR) for Somali, a currently extremely under-resourced language. This forms part of a continuing United Nations (UN) effort to employ ASR-based keyword spotting systems to support…

计算与语言 · 计算机科学 2019-07-09 Astik Biswas , Raghav Menon , Ewald van der Westhuizen , Thomas Niesler

Audio deepfake detection (ADD) is essential for preventing the misuse of synthetic voices that may infringe on personal rights and privacy. Recent zero-shot text-to-speech (TTS) models pose higher risks as they can clone voices with a…

声音 · 计算机科学 2024-09-23 Yuang Li , Min Zhang , Mengxin Ren , Miaomiao Ma , Daimeng Wei , Hao Yang

The development of Automatic Speech Recognition (ASR) systems for low-resource African languages remains challenging due to limited transcribed speech data. While recent advances in large multilingual models like OpenAI's Whisper offer…

计算与语言 · 计算机科学 2025-10-09 Benjamin Akera , Evelyn Nafula , Patrick Walukagga , Gilbert Yiga , John Quinn , Ernest Mwebaze

In recent years, datasets of paired audio and captions have enabled remarkable success in automatically generating descriptions for audio clips, namely Automated Audio Captioning (AAC). However, it is labor-intensive and time-consuming to…

声音 · 计算机科学 2023-09-22 Theodoros Kouzelis , Vassilis Katsouros

In this paper, we propose an end-to-end Mandarin tone classification method from continuous speech utterances utilizing both the spectrogram and the short-term context information as the input. Both spectrograms and context segment features…

声音 · 计算机科学 2021-12-20 Jiyang Tang , Ming Li

We propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework MMAudio. In contrast to single-modality training conditioned on (limited) video data…

计算机视觉与模式识别 · 计算机科学 2025-04-09 Ho Kei Cheng , Masato Ishii , Akio Hayakawa , Takashi Shibuya , Alexander Schwing , Yuki Mitsufuji

Sequence-to-sequence (seq2seq) models are competitive with hybrid models for automatic speech recognition (ASR) tasks when large amounts of training data are available. However, data sparsity and domain adaptation are more problematic for…

计算与语言 · 计算机科学 2021-06-16 Chak-Fai Li , Francis Keith , William Hartmann , Matthew Snover , Owen Kimball

Selecting in-domain data from a large pool of diverse and out-of-domain data is a non-trivial problem. In most cases simply using all of the available data will lead to sub-optimal and in some cases even worse performance compared to…

计算与语言 · 计算机科学 2019-07-03 Mortaza , Doulaty , Thomas Hain

Captioning has attracted much attention in image and video understanding while a small amount of work examines audio captioning. This paper contributes a Mandarin-annotated dataset for audio captioning within a car scene. A sentence-level…

声音 · 计算机科学 2020-10-26 Xuenan Xu , Heinrich Dinkel , Mengyue Wu , Kai Yu

The "MEG-MASC" dataset provides a curated set of raw magnetoencephalography (MEG) recordings of 27 English speakers who listened to two hours of naturalistic stories. Each participant performed two identical sessions, involving listening to…

定量方法 · 定量生物学 2022-08-25 Laura Gwilliams , Graham Flick , Alec Marantz , Liina Pylkkanen , David Poeppel , Jean-Remi King

Diverse promising datasets have been designed to hold back the development of fake audio detection, such as ASVspoof databases. However, previous datasets ignore an attacking situation, in which the hacker hides some small fake clips in…

声音 · 计算机科学 2023-12-19 Jiangyan Yi , Ye Bai , Jianhua Tao , Haoxin Ma , Zhengkun Tian , Chenglong Wang , Tao Wang , Ruibo Fu

The scarcity of large-scale classroom speech data has hindered the development of AI-driven speech models for education. Public classroom datasets remain limited, and the lack of a dedicated classroom noise corpus prevents the use of…

声音 · 计算机科学 2025-06-12 Ahmed Adel Attia , Jing Liu , Carl Espy-Wilson

Although automatic pathological speech detection approaches show promising results when clean recordings are available, they are vulnerable to additive noise. Recently it has been shown that databases commonly used to develop and evaluate…

音频与语音处理 · 电气工程与系统科学 2024-09-04 Mahdi Amiri , Ina Kodrasi

Wake word (WW) spotting is challenging in far-field not only because of the interference in signal transmission but also the complexity in acoustic environments. Traditional WW model training requires large amount of in-domain WW-specific…

音频与语音处理 · 电气工程与系统科学 2020-10-15 Yixin Gao , Yuriy Mishchenko , Anish Shah , Spyros Matsoukas , Shiv Vitaladevuni

Training neural text-to-speech (TTS) models for a new speaker typically requires several hours of high quality speech data. Prior works on voice cloning attempt to address this challenge by adapting pre-trained multi-speaker TTS models for…

声音 · 计算机科学 2022-04-07 Paarth Neekhara , Jason Li , Boris Ginsburg

Training a semi-supervised end-to-end speech recognition system using noisy student training has significantly improved performance. However, this approach requires a substantial amount of paired speech-text and unlabeled speech, which is…

计算与语言 · 计算机科学 2024-08-01 Chia-Yu Li , Ngoc Thang Vu

Achieving super-human performance in recognizing human speech has been a goal for several decades, as researchers have worked on increasingly challenging tasks. In the 1990's it was discovered, that conversational speech between two humans…

计算机视觉与模式识别 · 计算机科学 2021-07-28 Thai-Son Nguyen , Sebastian Stueker , Alex Waibel

Despite Telugu being spoken by over 80 million people, speech translation research for this morphologically rich language remains severely underexplored. We address this gap by developing a high-quality Telugu--English speech translation…

We propose a self-refining framework that enhances ASR performance with only unlabeled datasets. The process starts with an existing ASR model generating pseudo-labels on unannotated speech, which are then used to train a high-fidelity…