中文
相关论文

相关论文: Joint Fullband-Subband Modeling for High-Resolutio…

200 篇论文

Speech deepfake detection (SDD) systems perform well on standard benchmarks datasets but often fail to generalize to expressive and emotional spoofing attacks. Many methods rely on spoof-heavy training data, learning dataset-specific…

音频与语音处理 · 电气工程与系统科学 2026-04-16 Aurosweta Mahapatra , Ismail Rasim Ulgen , Kong Aik Lee , Nicholas Andrews , Berrak Sisman

Spectrograms - time-frequency representations of audio signals - have found widespread use in neural network-based spoofing detection. While deep models are trained on the fullband spectrum of the signal, we argue that not all frequency…

音频与语音处理 · 电气工程与系统科学 2020-04-07 Bhusan Chettri , Tomi Kinnunen , Emmanouil Benetos

Speech deepfake detection (SDD) is essential for maintaining trust in voice-driven technologies and digital media. Although recent SDD systems increasingly rely on self-supervised learning (SSL) representations that capture rich contextual…

音频与语音处理 · 电气工程与系统科学 2026-03-05 Cemal Hanilçi , Md Sahidullah , Tomi Kinnunen

Self-supervised representations excel at many vision and speech tasks, but their potential for audio-visual deepfake detection remains underexplored. Unlike prior work that uses these features in isolation or buried within complex…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Dragos-Alexandru Boldisor , Stefan Smeu , Dan Oneata , Elisabeta Oneata

Detecting singing-voice in polyphonic instrumental music is critical to music information retrieval. To train a robust vocal detector, a large dataset marked with vocal or non-vocal label at frame-level is essential. However, frame-level…

音频与语音处理 · 电气工程与系统科学 2020-08-12 Yuanbo Hou , Frank K. Soong , Jian Luan , Shengchen Li

This paper introduces an unsupervised framework for detecting audio patterns in musical samples (loops) through anomaly detection techniques, addressing challenges in music information retrieval (MIR). Existing methods are often constrained…

声音 · 计算机科学 2025-06-02 Shayan Dadman , Bernt Arild Bremdal , Børre Bang , Rune Dalmo

This paper proposes a full-band and sub-band fusion model, named as FullSubNet, for single-channel real-time speech enhancement. Full-band and sub-band refer to the models that input full-band and sub-band noisy spectral feature, output…

音频与语音处理 · 电气工程与系统科学 2024-07-04 Xiang Hao , Xiangdong Su , Radu Horaud , Xiaofei Li

Singing voice synthesis (SVS) systems are built to synthesize high-quality and expressive singing voice, in which the acoustic model generates the acoustic features (e.g., mel-spectrogram) given a music score. Previous singing acoustic…

音频与语音处理 · 电气工程与系统科学 2022-03-23 Jinglin Liu , Chengxi Li , Yi Ren , Feiyang Chen , Zhou Zhao

In this paper, we propose a deep-learning framework for environmental sound deepfake detection (ESDD) -- the task of identifying whether the sound scene and sound event in an input audio recording is fake or not. To this end, we conducted…

声音 · 计算机科学 2026-05-04 Lam Pham , Khoi Vu , Dat Tran , Phat Lam , Vu Nguyen , David Fischinger , Son Le

Significant strides have been made in creating voice identity representations using speech data. However, the same level of progress has not been achieved for singing voices. To bridge this gap, we suggest a framework for training singer…

声音 · 计算机科学 2024-01-11 Bernardo Torres , Stefan Lattner , Gaël Richard

Audio recorded in real-world environments often contains a mixture of foreground speech and background environmental sounds. With rapid advances in text-to-speech, voice conversion, and other generation models, either component can now be…

声音 · 计算机科学 2026-02-06 Xueping Zhang , Han Yin , Yang Xiao , Lin Zhang , Ting Dang , Rohan Kumar Das , Ming Li

Recent progress in audio generation models has made it possible to create highly realistic and immersive soundscapes, which are now widely used in film and virtual-reality-related applications. However, these audio generators also raise…

声音 · 计算机科学 2026-01-01 Han Yin , Yang Xiao , Rohan Kumar Das , Jisheng Bai , Ting Dang

Speech deepfake detection has achieved remarkable success in clean environments but faces significant challenges in complex, real-world scenarios where speech is often mixed with background music or noise. Current state-of-the-art methods…

声音 · 计算机科学 2026-05-25 Qingcao Li , Yipeng Lin , Weichen Lian , Zhongjie Ba , Peng Cheng , Zhichao Lian

Singing Voice Conversion (SVC) is a technique that enables any singer to perform any song. To achieve this, it is essential to obtain speaker-agnostic representations from the source audio, which poses a significant challenge. A common…

声音 · 计算机科学 2024-09-17 Xueyao Zhang , Zihao Fang , Yicheng Gu , Haopeng Chen , Lexiao Zou , Junan Zhang , Liumeng Xue , Zhizheng Wu

Audio deepfake detection (ADD) is crucial to combat the misuse of speech synthesized from generative AI models. Existing ADD models suffer from generalization issues, with a large performance discrepancy between in-domain and out-of-domain…

声音 · 计算机科学 2024-07-29 Yi Zhu , Surya Koppisetti , Trang Tran , Gaurav Bharaj

This paper describes a submission to the Environment-Aware Speech and Sound Deepfake Detection Challenge (ESDD2) 2026, which addresses component-level deepfake detection using the CompSpoofV2 dataset, where speech and environmental sounds…

声音 · 计算机科学 2026-05-06 Khalid Zaman , Qixuan Huang , Muhammad Uzair , Masashi Unoki

Recent advances in synthetic speech have made audio deepfakes increasingly realistic, posing significant security risks. Existing detection methods that rely on a single modality, either raw waveform embeddings or spectral based features,…

In this paper, we develop DeepSinger, a multi-lingual multi-singer singing voice synthesis (SVS) system, which is built from scratch using singing training data mined from music websites. The pipeline of DeepSinger consists of several…

音频与语音处理 · 电气工程与系统科学 2020-07-16 Yi Ren , Xu Tan , Tao Qin , Jian Luan , Zhou Zhao , Tie-Yan Liu

The rapid advancement of AI has enabled highly realistic speech synthesis and voice cloning, posing serious risks to voice authentication, smart assistants, and telecom security. While most prior work frames spoof detection as a binary…

声音 · 计算机科学 2025-09-10 Bin Hu , Kunyang Huang , Daehan Kwak , Meng Xu , Kuan Huang

This paper proposes singing voice synthesis (SVS) based on frame-level sequence-to-sequence models considering vocal timing deviation. In SVS, it is essential to synchronize the timing of singing with temporal structures represented by…

音频与语音处理 · 电气工程与系统科学 2023-02-23 Miku Nishihara , Yukiya Hono , Kei Hashimoto , Yoshihiko Nankaku , Keiichi Tokuda