中文
相关论文

相关论文: ASDKit: A Toolkit for Comprehensive Evaluation of …

200 篇论文

Recently, more and more personalized speech enhancement systems (PSE) with excellent performance have been proposed. However, two critical issues still limit the performance and generalization ability of the model: 1) Acoustic environment…

音频与语音处理 · 电气工程与系统科学 2022-11-23 Xiaofeng Ge , Jiangyu Han , Haixin Guan , Yanhua Long

Recent advances in the Active Speaker Detection (ASD) problem build upon a two-stage process: feature extraction and spatio-temporal context aggregation. In this paper, we propose an end-to-end ASD workflow where feature learning and…

计算机视觉与模式识别 · 计算机科学 2022-07-26 Juan Leon Alcazar , Moritz Cordes , Chen Zhao , Bernard Ghanem

We propose an algorithm for detecting patterns exhibited by anomalous clusters in high dimensional discrete data. Unlike most anomaly detection (AD) methods, which detect individual anomalies, our proposed method detects groups (clusters)…

机器学习 · 统计学 2016-05-23 Hossein Soleimani , David J. Miller

In this paper, we propose a deep-learning framework for environmental sound deepfake detection (ESDD) -- the task of identifying whether the sound scene and sound event in an input audio recording is fake or not. To this end, we conducted…

声音 · 计算机科学 2026-05-04 Lam Pham , Khoi Vu , Dat Tran , Phat Lam , Vu Nguyen , David Fischinger , Son Le

In this paper, we propose a new strategy for acoustic scene classification (ASC) , namely recognizing acoustic scenes through identifying distinct sound events. This differs from existing strategies, which focus on characterizing global…

声音 · 计算机科学 2019-10-23 Hongwei Song , Jiqing Han , Shiwen Deng , Zhihao Du

Recognition and interpretation of bird vocalizations are pivotal in ornithological research and ecological conservation efforts due to their significance in understanding avian behaviour, performing habitat assessment and judging ecological…

音频与语音处理 · 电气工程与系统科学 2024-07-30 Yashwardhan Chaudhuri , Paridhi Mundra , Arnesh Batra , Orchid Chetia Phukan , Arun Balaji Buduru

Fake audio detection is a growing concern and some relevant datasets have been designed for research. However, there is no standard public Chinese dataset under complex conditions.In this paper, we aim to fill in the gap and design a…

声音 · 计算机科学 2023-07-19 Haoxin Ma , Jiangyan Yi , Chenglong Wang , Xinrui Yan , Jianhua Tao , Tao Wang , Shiming Wang , Ruibo Fu

In this paper, we present a new open source toolkit for automatic speech recognition (ASR), named CAT (CRF-based ASR Toolkit). A key feature of CAT is discriminative training in the framework of conditional random field (CRF), particularly…

机器学习 · 计算机科学 2019-11-21 Keyu An , Hongyu Xiang , Zhijian Ou

In this paper, we propose a deep learning based model for Acoustic Anomaly Detection of Machines, the task for detecting abnormal machines by analysing the machine sound. By conducting extensive experiments, we indicate that multiple…

音频与语音处理 · 电气工程与系统科学 2024-03-04 Tin Nguyen , Lam Pham , Phat Lam , Dat Ngo , Hieu Tang , Alexander Schindler

Many datasets have been designed to further the development of fake audio detection. However, fake utterances in previous datasets are mostly generated by altering timbre, prosody, linguistic content or channel noise of original audio.…

Auditory attention decoding (AAD) is the process of identifying the attended speech in a multi-talker environment using brain signals, typically recorded through electroencephalography (EEG). Over the past decade, AAD has undergone…

声音 · 计算机科学 2025-07-08 Nhan Duc Thanh Nguyen , Huy Phan , Simon Geirnaert , Kaare Mikkelsen , Preben Kidmose

The availability of audio data on sound sharing platforms such as Freesound gives users access to large amounts of annotated audio. Utilising such data for training is becoming increasingly popular, but the problem of label noise that is…

声音 · 计算机科学 2022-03-01 Turab Iqbal , Yin Cao , Andrew Bailey , Mark D. Plumbley , Wenwu Wang

We introduce a toolkit for uncovering spurious correlations between recording characteristics and target class in speech datasets. Spurious correlations may arise due to heterogeneous recording conditions, a common scenario for…

Audio Deepfake Detection (ADD) aims to detect spoof speech from bonafide speech. Most prior studies assume that stronger correlations within or across acoustic and emotional features imply authenticity, and thus focus on enhancing or…

声音 · 计算机科学 2026-01-21 Jinhua Zhang , Zhenqi Jia , Rui Liu

Audio deepfake detection (ADD) is crucial to combat the misuse of speech synthesized from generative AI models. Existing ADD models suffer from generalization issues, with a large performance discrepancy between in-domain and out-of-domain…

声音 · 计算机科学 2024-07-29 Yi Zhu , Surya Koppisetti , Trang Tran , Gaurav Bharaj

Multimedia anomaly datasets play a crucial role in automated surveillance. They have a wide range of applications expanding from outlier objects/ situation detection to the detection of life-threatening events. For more than 1.5 decades,…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Pratibha Kumari , Anterpreet Kaur Bedi , Mukesh Saini

This research presents Muskits-ESPnet, a versatile toolkit that introduces new paradigms to Singing Voice Synthesis (SVS) through the application of pretrained audio models in both continuous and discrete approaches. Specifically, we…

Different machines can exhibit diverse frequency patterns in their emitted sound. This feature has been recently explored in anomaly sound detection and reached state-of-the-art performance. However, existing methods rely on the manual or…

声音 · 计算机科学 2023-09-07 Hejing Zhang , Jian Guan , Qiaoxi Zhu , Feiyang Xiao , Youde Liu

Constructing a dataset for replay spoofing detection requires a physical process of playing an utterance and re-recording it, presenting a challenge to the collection of large-scale datasets. In this study, we propose a self-supervised…

机器学习 · 计算机科学 2020-08-20 Hye-jin Shim , Hee-Soo Heo , Jee-weon Jung , Ha-Jin Yu

Audio deepfake detection is an emerging active topic. A growing number of literatures have aimed to study deepfake detection algorithms and achieved effective performance, the problem of which is far from being solved. Although there are…

声音 · 计算机科学 2023-08-30 Jiangyan Yi , Chenglong Wang , Jianhua Tao , Xiaohui Zhang , Chu Yuan Zhang , Yan Zhao