English
Related papers

Related papers: FMFCC-A: A Challenging Mandarin Dataset for Synthe…

200 papers

Whisper speech recognition is crucial not only for ensuring privacy in sensitive communications but also for providing a critical communication bridge for patients under vocal restraint and enabling discrete interaction in noise-sensitive…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 Cancan Li , Fei Su , Juan Liu , Hui Bu , Yulong Wan , Hongbin Suo , Ming Li

Recent state-of-the-art neural text-to-speech (TTS) synthesis models have dramatically improved intelligibility and naturalness of generated speech from text. However, building a good bilingual or code-switched TTS for a particular voice is…

Sound · Computer Science 2020-10-19 Shengkui Zhao , Trung Hieu Nguyen , Hao Wang , Bin Ma

With the rapid advancement of technologies like text-to-speech (TTS) and voice conversion (VC), detecting deepfake voices has become increasingly crucial. However, both academia and industry lack a comprehensive and intuitive benchmark for…

Sound · Computer Science 2024-09-11 Ziwei Yan , Yanjie Zhao , Haoyu Wang

Aphasia is a language disorder that affects the speaking ability of millions of patients. This paper presents a new benchmark for Aphasia speech recognition and detection tasks using state-of-the-art speech recognition techniques with the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-24 Jiyang Tang , William Chen , Xuankai Chang , Shinji Watanabe , Brian MacWhinney

This paper introduces a high-quality open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide. The dataset consists of about 93 hours of transcribed audio recordings spoken by two…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-09 Saida Mussakhojayeva , Aigerim Janaliyeva , Almas Mirzakhmetov , Yerbolat Khassanov , Huseyin Atakan Varol

The rapid advancement of generative models has enabled highly realistic audio deepfakes, yet current detectors suffer from a critical bias problem, leading to poor generalization across unseen datasets. This paper proposes Artifact-Focused…

Cued Speech (CS) is a communication system developed for deaf people, which exploits hand cues to complement speechreading at the phonetic level. Currently, it is estimated that CS has been adapted to over 60 languages; however, no official…

Audio and Speech Processing · Electrical Eng. & Systems 2020-01-06 Liu Li , Feng Gang

Recent advances in speech synthesis and voice conversion have greatly improved the naturalness and authenticity of generated audio. Meanwhile, evolving encoding, compression, and transmission mechanisms on social media platforms further…

Sound · Computer Science 2026-03-09 Daixian Li , Jun Xue , Yanzhen Ren , Zhuolin Yi , Yihuan Huang , Guanxiang Feng , Yi Chai

Active speaker detection is an important component in video analysis algorithms for applications such as speaker diarization, video re-targeting for meetings, speech enhancement, and human-robot interaction. The absence of a large,…

Generative AI advances rapidly, allowing the creation of very realistic manipulated video and audio. This progress presents a significant security and ethical threat, as malicious users can exploit DeepFake techniques to spread…

Multimedia · Computer Science 2025-06-09 Marcel Klemt , Carlotta Segna , Anna Rohrbach

With the development of large text-to-speech (TTS) models and scale-up of the training data, state-of-the-art TTS systems have achieved impressive performance. In this paper, we present WenetSpeech4TTS, a multi-domain Mandarin corpus…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-21 Linhan Ma , Dake Guo , Kun Song , Yuepeng Jiang , Shuai Wang , Liumeng Xue , Weiming Xu , Huan Zhao , Binbin Zhang , Lei Xie

The development of high-performance, on-device keyword spotting (KWS) systems for ultra-low-power hardware is critically constrained by the scarcity of specialized, multi-command training datasets. Traditional data collection through human…

Sound · Computer Science 2025-11-25 Lu Gan , Xi Li

Humans can perceive speakers' characteristics (e.g., identity, gender, personality and emotion) by their appearance, which are generally aligned to their voice style. Recently, vision-driven Text-to-speech (TTS) scholars grounded their…

Sound · Computer Science 2025-04-17 Tian-Hao Zhang , Jiawei Zhang , Jun Wang , Xinyuan Qian , Xu-Cheng Yin

Recent diffusion-based text-to-speech (TTS) models achieve high naturalness and expressiveness, yet often suffer from speaker drift, a subtle, gradual shift in perceived speaker identity within a single utterance. This underexplored…

Machine Reading Comprehension (MRC) has become enormously popular recently and has attracted a lot of attention. However, the existing reading comprehension datasets are mostly in English. In this paper, we introduce a Span-Extraction…

Computation and Language · Computer Science 2019-11-05 Yiming Cui , Ting Liu , Wanxiang Che , Li Xiao , Zhipeng Chen , Wentao Ma , Shijin Wang , Guoping Hu

AI-synthesized speech, also known as deepfake speech, has recently raised significant concerns due to the rapid advancement of speech synthesis and speech conversion techniques. Previous works often rely on distinguishing synthesizer…

Sound · Computer Science 2024-11-15 Kuiyuan Zhang , Zhongyun Hua , Yushu Zhang , Yifang Guo , Tao Xiang

The recent advancements in generative artificial speech models have made possible the generation of highly realistic speech signals. At first, it seems exciting to obtain these artificially synthesized signals such as speech clones or deep…

Sound · Computer Science 2022-03-09 Karan Bhatia , Ansh Agrawal , Priyanka Singh , Arun Kumar Singh

Despite having hundreds of millions of speakers, Chinese dialects lag behind Mandarin in speech and language technologies. Most varieties are primarily spoken, making dialect-to-Mandarin speech-LLMs (large language models) more practical…

Computation and Language · Computer Science 2026-01-13 Kalvin Chang , Yiwen Shao , Jiahong Li , Dong Yu

The existing audio datasets are predominantly tailored towards single languages, overlooking the complex linguistic behaviors of multilingual communities that engage in code-switching. This practice, where individuals frequently mix two or…

Sound · Computer Science 2025-03-04 Peng Xie , Kani Chen

Chinese Grammatical Error Correction (CGEC) has been attracting growing attention from researchers recently. In spite of the fact that multiple CGEC datasets have been developed to support the research, these datasets lack the ability to…

Computation and Language · Computer Science 2023-11-10 Hanyue Du , Yike Zhao , Qingyuan Tian , Jiani Wang , Lei Wang , Yunshi Lan , Xuesong Lu
‹ Prev 1 3 4 5 6 7 10 Next ›