中文
相关论文

相关论文: An Empirical Analysis of Task-Induced Encoder Bias…

200 篇论文

We propose the Fr\'echet Audio Distance (FAD), a novel, reference-free evaluation metric for music enhancement algorithms. We demonstrate how typical evaluation metrics for speech enhancement and blind source separation can fail to…

音频与语音处理 · 电气工程与系统科学 2019-01-18 Kevin Kilgour , Mauricio Zuluaga , Dominik Roblek , Matthew Sharifi

The growing popularity of generative music models underlines the need for perceptually relevant, objective music quality metrics. The Frechet Audio Distance (FAD) is commonly used for this purpose even though its correlation with perceptual…

音频与语音处理 · 电气工程与系统科学 2024-03-07 Azalea Gui , Hannes Gamper , Sebastian Braun , Dimitra Emmanouilidou

This paper explores whether considering alternative domain-specific embeddings to calculate the Fr\'echet Audio Distance (FAD) metric can help the FAD to correlate better with perceptual ratings of environmental sounds. We used embeddings…

The complex nature of musical emotion introduces inherent bias in both recognition and generation, particularly when relying on a single audio encoder, emotion classifier, or evaluation metric. In this work, we conduct a study on Music…

音频与语音处理 · 电气工程与系统科学 2025-05-01 Yuanchao Li , Azalea Gui , Dimitra Emmanouilidou , Hannes Gamper

Although being widely adopted for evaluating generated audio signals, the Fr\'echet Audio Distance (FAD) suffers from significant limitations, including reliance on Gaussian assumptions, sensitivity to sample size, and high computational…

声音 · 计算机科学 2025-03-11 Yoonjin Chung , Pilsun Eu , Junwon Lee , Keunwoo Choi , Juhan Nam , Ben Sangbae Chon

Despite significant recent advances in generative acoustic text-to-music (TTM) modeling, robust evaluation of these models lags behind, relying in particular on the popular Fr\'echet Audio Distance (FAD). In this work, we rigorously study…

Neural audio codecs (NACs) achieve low-bitrate compression by learning compact audio representations, which can also serve as features for perceptual quality evaluation. We introduce DACe, an enhanced, higher-fidelity version of the…

音频与语音处理 · 电气工程与系统科学 2026-01-23 Arijit Biswas , Lars Villemoes

Automated audio captioning aims to use natural language to describe the content of audio data. This paper presents an audio captioning system with an encoder-decoder architecture, where the decoder predicts words based on audio features…

音频与语音处理 · 电气工程与系统科学 2021-08-06 Xinhao Mei , Qiushi Huang , Xubo Liu , Gengyun Chen , Jingqian Wu , Yusong Wu , Jinzheng Zhao , Shengchen Li , Tom Ko , H Lilian Tang , Xi Shao , Mark D. Plumbley , Wenwu Wang

Blind Estimation of Audio Effects (BE-AFX) aims at estimating the Audio Effects (AFXs) applied to an original, unprocessed audio sample solely based on the processed audio sample. To train such a system traditional approaches optimize a…

声音 · 计算机科学 2024-02-12 Côme Peladeau , Geoffroy Peeters

Embedding acoustic information into fixed length representations is of interest for a whole range of applications in speech and audio technology. Two novel unsupervised approaches to generate acoustic embeddings by modelling of acoustic…

计算与语言 · 计算机科学 2021-02-08 Yanpei Shi , Thomas Hain

One way to interpret trained deep neural networks (DNNs) is by inspecting characteristics that neurons in the model respond to, such as by iteratively optimising the model input (e.g., an image) to maximally activate specific neurons.…

机器学习 · 计算机科学 2019-07-02 Saumitra Mishra , Daniel Stoller , Emmanouil Benetos , Bob L. Sturm , Simon Dixon

Human listeners exhibit the remarkable ability to segregate a desired sound from complex acoustic scenes through selective auditory attention, motivating the study of Targeted Sound Detection (TSD). The task requires detecting and…

音频与语音处理 · 电气工程与系统科学 2026-03-19 Shubham Gupta , Adarsh Arigala , B. R. Dilleswari , Sri Rama Murty Kodukula

A great interest has arisen in using Deep Generative Models (DGM) for generative design. When assessing the quality of the generated designs, human designers focus more on structural plausibility, e.g., no missing component, rather than…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Jiajie Fan , Amal Trigui , Thomas Bäck , Hao Wang

We consider the problem of task-agnostic feature upsampling in dense prediction where an upsampling operator is required to facilitate both region-sensitive tasks like semantic segmentation and detail-sensitive tasks such as image matting.…

计算机视觉与模式识别 · 计算机科学 2022-12-29 Hao Lu , Wenze Liu , Hongtao Fu , Zhiguo Cao

The generalization of Fake Audio Detection (FAD) is critical due to the emergence of new spoofing techniques. Traditional FAD methods often focus solely on distinguishing between genuine and known spoofed audio. We propose a Genuine-Focused…

Recent promising results in auditory attention decoding (AAD) using scalp electroencephalography (EEG) have motivated the exploration of cEEGrid, a flexible and portable ear-EEG system. While prior cEEGrid-based studies have confirmed the…

音频与语音处理 · 电气工程与系统科学 2025-10-23 Yuanming Zhang , Zeyan Song , Jing Lu , Fei Chen , Zhibin Lin

Anomalous sound detection (ASD) typically involves self-supervised proxy tasks to learn feature representations from normal sound data, owing to the scarcity of anomalous samples. In ASD research, proxy tasks such as AutoEncoders operate…

音频与语音处理 · 电气工程与系统科学 2026-01-14 Seunghyeon Shin , Seokjin Lee

Speaker embeddings become growing popular in the text-independent speaker verification task. In this paper, we propose two improvements during the training stage. The improvements are both based on triplet cause the training stage and the…

音频与语音处理 · 电气工程与系统科学 2019-08-08 Zongze Ren , Zhiyong Chen , Shugong Xu

Evaluating text-to-image and text-to-video models is challenging due to a fundamental disconnect: established metrics fail to jointly measure visual quality and semantic alignment with text, leading to a poor correlation with human…

计算机视觉与模式识别 · 计算机科学 2025-08-28 Jaywon Koo , Jefferson Hernandez , Moayed Haji-Ali , Ziyan Yang , Vicente Ordonez

Automatic Mean Opinion Score (MOS) prediction is employed to evaluate the quality of synthetic speech. This study extends the application of predicted MOS to the task of Fake Audio Detection (FAD), as we expect that MOS can be used to…

音频与语音处理 · 电气工程与系统科学 2024-01-26 Wangjin Zhou , Zhengdong Yang , Chenhui Chu , Sheng Li , Raj Dabre , Yi Zhao , Tatsuya Kawahara
‹ 上一页 1 2 3 10 下一页 ›