English
Related papers

Related papers: A Semantically Consistent Dataset for Data-Efficie…

200 papers

The rapid advancement of AI has enabled highly realistic speech synthesis and voice cloning, posing serious risks to voice authentication, smart assistants, and telecom security. While most prior work frames spoof detection as a binary…

Sound · Computer Science 2025-09-10 Bin Hu , Kunyang Huang , Daehan Kwak , Meng Xu , Kuan Huang

Segmenting audio into homogeneous sections such as music and speech helps us understand the content of audio. It is useful as a pre-processing step to index, store, and modify audio recordings, radio broadcasts and TV programmes. Deep…

Audio fingerprinting systems must efficiently and robustly identify query snippets in an extensive database. To this end, state-of-the-art systems use deep learning to generate compact audio fingerprints. These systems deploy indexing…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-20 Anup Singh , Kris Demuynck , Vipul Arora

Recent years have seen progress beyond domain-specific sound separation for speech or music towards universal sound separation for arbitrary sounds. Prior work on universal sound separation has investigated separating a target sound out of…

Dialogue state tracking plays a crucial role in extracting information in task-oriented dialogue systems. However, preceding research are limited to textual modalities, primarily due to the shortage of authentic human audio datasets. We…

Sound · Computer Science 2023-12-05 Jihyun Lee , Yejin Jeon , Wonjun Lee , Yunsu Kim , Gary Geunbae Lee

While significant advances have been made with respect to the separation of overlapping speech signals, studies have been largely constrained to mixtures of clean, near anechoic speech, not representative of many real-world scenarios.…

Sound · Computer Science 2020-02-17 Matthew Maciejewski , Gordon Wichern , Emmett McQuinn , Jonathan Le Roux

In this paper, we propose a method called Hodge and Podge for sound event detection. We demonstrate Hodge and Podge on the dataset of Detection and Classification of Acoustic Scenes and Events (DCASE) 2019 Challenge Task 4. This task aims…

Sound · Computer Science 2020-02-17 Ziqiang Shi , Liu Liu , Huibin Lin , Rujie Liu

Whisper generation is constrained by the difficulty of data collection. Because whispered speech has low acoustic amplitude, high-fidelity recording is challenging. In this paper, we introduce WhispSynth, a large-scale multilingual corpus…

Sound · Computer Science 2026-03-17 Tianyi Tan , Jiaxin Ye , Yuanming Zhang , Xiaohuai Le , Xianjun Xia , Chuanzeng Huang , Jing Lu

Our goal is to collect a large-scale audio-visual dataset with low label noise from videos in the wild using computer vision techniques. The resulting dataset can be used for training and evaluating audio recognition models. We make three…

Computer Vision and Pattern Recognition · Computer Science 2020-09-28 Honglie Chen , Weidi Xie , Andrea Vedaldi , Andrew Zisserman

As several large single-dish radio surveys begin operation within the coming decade, a wealth of radio data will become available and provide a new window to the Universe. In order to fully exploit the potential of these data sets, it is…

Instrumentation and Methods for Astrophysics · Physics 2017-01-16 Joel Akeret , Sebastian Seehars , Chihway Chang , Christian Monstein , Adam Amara , Alexandre Refregier

Distance estimation from audio plays a crucial role in various applications, such as acoustic scene analysis, sound source localization, and room modeling. Most studies predominantly center on employing a classification approach, where…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-27 Michael Neri , Archontis Politis , Daniel Krause , Marco Carli , Tuomas Virtanen

Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings. The association of these constituent sound events with their mixture and…

This work presents a supervised deep hashing method for retrieving similar audio events. The proposed method, named AudioNet, is a deep-learning-based system for efficient hashing and retrieval of similar audio events using an audio example…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-04 Sagar Dutta , Vipul Arora

Modern large language models often encode sensitive, harmful, or copyrighted knowledge, raising the need for post-hoc unlearning-the ability to remove specific domains of knowledge from a model without full retraining. A major bottleneck in…

Computation and Language · Computer Science 2025-10-08 Xiaoyuan Zhu , Muru Zhang , Ollie Liu , Robin Jia , Willie Neiswanger

We propose an automatic data processing pipeline to extract vocal productions from large-scale natural audio recordings and classify these vocal productions. The pipeline is based on a deep neural network and adresses both issues…

Audio recorded in real-world environments often contains a mixture of foreground speech and background environmental sounds. With rapid advances in text-to-speech, voice conversion, and other generation models, either component can now be…

Sound · Computer Science 2026-02-06 Xueping Zhang , Han Yin , Yang Xiao , Lin Zhang , Ting Dang , Rohan Kumar Das , Ming Li

Speech synthesis systems can now produce highly realistic vocalisations that pose significant authenticity challenges. Despite substantial progress in deepfake detection models, their real-world effectiveness is often undermined by evolving…

Sound · Computer Science 2026-02-12 Qizhou Wang , Hanxun Huang , Guansong Pang , Sarah Erfani , Christopher Leckie

Speech recognition in highly-reverberant real environments remains a major challenge. An evaluation dataset for this task is needed. This report describes the generation of the Highly-Reverberant Real Environment database (HRRE). This…

Audio and Speech Processing · Electrical Eng. & Systems 2018-03-28 Juan Pablo Escudero , Victor Poblete , José Novoa , Jorge Wuth , Josué Fredes , Rodrigo Mahu , Richard Stern , Néstor Becerra Yoma

State-of-the-art under-determined audio source separation systems rely on supervised end-end training of carefully tailored neural network architectures operating either in the time or the spectral domain. However, these methods are…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-29 Vivek Narayanaswamy , Jayaraman J. Thiagarajan , Rushil Anirudh , Andreas Spanias

Component-level audio Spoofing (Comp-Spoof) targets a new form of audio manipulation where only specific components of a signal, such as speech or environmental sound, are forged or substituted while other components remain genuine.…

Sound · Computer Science 2026-02-02 Xueping Zhang , Yechen Wang , Linxi Li , Liwei Jin , Ming Li