中文
相关论文

相关论文: FMFCC-A: A Challenging Mandarin Dataset for Synthe…

200 篇论文

With the rise of low power speech-enabled devices, there is a growing demand to quickly produce models for recognizing arbitrary sets of keywords. As with many machine learning tasks, one of the most challenging parts in the model creation…

音频与语音处理 · 电气工程与系统科学 2020-02-05 James Lin , Kevin Kilgour , Dominik Roblek , Matthew Sharifi

Dialogue state tracking plays a crucial role in extracting information in task-oriented dialogue systems. However, preceding research are limited to textual modalities, primarily due to the shortage of authentic human audio datasets. We…

声音 · 计算机科学 2023-12-05 Jihyun Lee , Yejin Jeon , Wonjun Lee , Yunsu Kim , Gary Geunbae Lee

The rapid surge of text-to-speech and face-voice reenactment models makes video fabrication easier and highly realistic. To encounter this problem, we require datasets that rich in type of generation methods and perturbation strategy which…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Zhixi Cai , Kartik Kuckreja , Shreya Ghosh , Akanksha Chuchra , Muhammad Haris Khan , Usman Tariq , Tom Gedeon , Abhinav Dhall

Current state of the art acoustic models can easily comprise more than 100 million parameters. This growing complexity demands larger training datasets to maintain a decent generalization of the final decision function. An ideal dataset is…

音频与语音处理 · 电气工程与系统科学 2022-02-01 Philipp Klumpp , Tomás Arias-Vergara , Paula Andrea Pérez-Toro , Elmar Nöth , Juan Rafael Orozco-Arroyave

This paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in multi-channel format for diarization and automatic speech…

音频与语音处理 · 电气工程与系统科学 2022-02-07 Naijun Zheng , Na Li , Xixin Wu , Lingwei Meng , Jiawen Kang , Haibin Wu , Chao Weng , Dan Su , Helen Meng

In this paper, we present AISHELL-3, a large-scale and high-fidelity multi-speaker Mandarin speech corpus which could be used to train multi-speaker Text-to-Speech (TTS) systems. The corpus contains roughly 85 hours of emotion-neutral…

声音 · 计算机科学 2021-04-23 Yao Shi , Hui Bu , Xin Xu , Shaoji Zhang , Ming Li

We present the first large-scale open-set benchmark for multilingual audio-video deepfake detection. Our dataset comprises over 250 hours of real and fake videos across eight languages, with 60% of data being generated. For each language,…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Florinel-Alin Croitoru , Vlad Hondru , Marius Popescu , Radu Tudor Ionescu , Fahad Shahbaz Khan , Mubarak Shah

Despite extensive research on textual and visual disambiguation, disambiguation through speech (DTS) remains underexplored. This is largely due to the lack of high-quality datasets that pair spoken sentences with richly ambiguous text. To…

计算与语言 · 计算机科学 2025-06-10 Haotian Guo , Jing Han , Yongfeng Tu , Shihao Gao , Shengfan Shen , Wulong Xiang , Weihao Gan , Zixing Zhang

In this technical report, we describe our submission for the WildSpoof Challenge TTS Track: Text-to-Speech with In-the-Wild Data. We introduce F5-TTS-DPS, a model built upon the F5-TTS architecture. Our approach integrates Exponential…

音频与语音处理 · 电气工程与系统科学 2026-05-25 Renhe Sun , Jiayi Zhou , Haolin He , Yueying Feng , Jian Liu

In this paper, we present CSCD-NS, the first Chinese spelling check (CSC) dataset designed for native speakers, containing 40,000 samples from a Chinese social platform. Compared with existing CSC datasets aimed at Chinese learners, CSCD-NS…

计算与语言 · 计算机科学 2024-05-24 Yong Hu , Fandong Meng , Jie Zhou

The clinical diagnosis of most mental disorders primarily relies on the conversations between psychiatrist and patient. The creation of such diagnostic conversation datasets is promising to boost the AI mental healthcare community. However,…

计算与语言 · 计算机科学 2024-12-30 Congchi Yin , Feng Li , Shu Zhang , Zike Wang , Jun Shao , Piji Li , Jianhua Chen , Xun Jiang

Existing fake audio detection systems perform well in in-domain testing, but still face many challenges in out-of-domain testing. This is due to the mismatch between the training and test data, as well as the poor generalizability of…

声音 · 计算机科学 2023-05-24 Chenglong Wang , Jiangyan Yi , Jianhua Tao , Chuyuan Zhang , Shuai Zhang , Xun Chen

Large-scale datasets have successively proven their fundamental importance in several research fields, especially for early progress in some emerging topics. In this paper, we focus on the problem of visual speech recognition, also known as…

计算机视觉与模式识别 · 计算机科学 2019-04-25 Shuang Yang , Yuanhang Zhang , Dalu Feng , Mingmin Yang , Chenhao Wang , Jingyun Xiao , Keyu Long , Shiguang Shan , Xilin Chen

Recent studies have proposed the use of Text-To-Speech (TTS) systems to automatically synthesise speech test cases on a scale and uncover a large number of failures in ASR systems. However, the failures uncovered by synthetic test cases may…

Synthesized speech is common today due to the prevalence of virtual assistants, easy-to-use tools for generating and modifying speech signals, and remote work practices. Synthesized speech can also be used for nefarious purposes, including…

声音 · 计算机科学 2022-05-05 Emily R. Bartusiak , Edward J. Delp

Advancements in artificial intelligence and machine learning have significantly improved synthetic speech generation. This paper explores diffusion models, a novel method for creating realistic synthetic speech. We create a diffusion…

密码学与安全 · 计算机科学 2025-01-15 Anton Firc , Kamil Malinka , Petr Hanáček

Traditional voice conversion (VC) methods typically attempt to separate speaker identity and linguistic information into distinct representations, which are then combined to reconstruct the audio. However, effectively disentangling these…

声音 · 计算机科学 2025-10-13 Huu Tuong Tu , Huan Vu , cuong tien nguyen , Dien Hy Ngo , Nguyen Thi Thu Trang

This paper introduces SpoofCeleb, a dataset designed for Speech Deepfake Detection (SDD) and Spoofing-robust Automatic Speaker Verification (SASV), utilizing source data from real-world conditions and spoofing attacks generated by…

The rapid advancement of generative models has enabled the creation of increasingly stealthy synthetic voices, commonly referred to as audio deepfakes. A recent technique, FOICE [USENIX'24], demonstrates a particularly alarming capability:…

密码学与安全 · 计算机科学 2025-11-14 Nguyen Linh Bao Nguyen , Alsharif Abuadbba , Kristen Moore , Tingmin Wu

The rise of deepfake audio and hate speech, powered by advanced text-to-speech, threatens online safety. We present SynHate, the first multilingual dataset for detecting hate speech in synthetic audio, spanning 37 languages. SynHate uses a…

声音 · 计算机科学 2025-06-10 Rishabh Ranjan , Kishan Pipariya , Mayank Vatsa , Richa Singh