English
Related papers

Related papers: Parameter Sensitivity of Deep-Feature based Evalua…

200 papers

We address the problem of estimating depth with multi modal audio visual data. Inspired by the ability of animals, such as bats and dolphins, to infer distance of objects with echolocation, some recent methods have utilized echoes for depth…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Kranti Kumar Parida , Siddharth Srivastava , Gaurav Sharma

In audio processing applications, phase retrieval (PR) is often performed from the magnitude of short-time Fourier transform (STFT) coefficients. Although PR performance has been observed to depend on the considered STFT parameters and…

Signal Processing · Electrical Eng. & Systems 2021-06-10 Andrés Marafioti , Nicki Holighaus , Piotr Majdak

Automated dysarthria detection and severity assessment from speech have attracted significant research attention due to their potential clinical impact. Despite rapid progress in acoustic modeling and deep learning, models still fall short…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-28 Krishna Gurugubelli

Audio embeddings are crucial tools in understanding large catalogs of music. Typically embeddings are evaluated on the basis of the performance they provide in a wide range of downstream tasks, however few studies have investigated the…

Current perceptual similarity metrics operate at the level of pixels and patches. These metrics compare images in terms of their low-level colors and textures, but fail to capture mid-level similarities and differences in image layout,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Stephanie Fu , Netanel Tamir , Shobhita Sundaram , Lucy Chai , Richard Zhang , Tali Dekel , Phillip Isola

Music similarity search is useful for a variety of creative tasks such as replacing one music recording with another recording with a similar "feel", a common task in video editing. For this task, it is typically necessary to define a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-14 Jongpil Lee , Nicholas J. Bryan , Justin Salamon , Zeyu Jin , Juhan Nam

In some DNNs for audio source separation, the relevant model parameters are independent of the sampling frequency of the audio used for training. Considering the application of dialogue separation, this is shown for two DNN architectures: a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-07 Jouni Paulus , Matteo Torcoli

Deep Metric Learning (DML) methods aim at learning an embedding space in which distances are closely related to the inherent semantic similarity of the inputs. Previous studies have shown that popular benchmark datasets often contain…

Computer Vision and Pattern Recognition · Computer Science 2023-11-02 Oriol Barbany , Xiaofan Lin , Muhammet Bastan , Arnab Dhua

Objective: Speech tests aim to estimate discrimination loss or speech recognition threshold (SRT). This paper investigates the potential to estimate SRTs from clinical data that target at characterizing the discrimination loss. Knowledge…

Sound · Computer Science 2025-01-16 Mareike Buhl , Eugen Kludt , Lena Schell-Majoor , Paul Avan , Marta Campi

Automatic Music Transcription (AMT) -- the task of converting music audio into note representations -- has seen rapid progress, driven largely by deep learning systems. Due to the limited availability of richly annotated music datasets,…

Sound · Computer Science 2026-01-27 Lukáš Samuel Marták , Patricia Hu , Gerhard Widmer

Several methods have recently been proposed to analyze speech and automatically infer the personality of the speaker. These methods often rely on prosodic and other hand crafted speech processing features extracted with off-the-shelf…

Computer Vision and Pattern Recognition · Computer Science 2022-05-10 Marc-André Carbonneau , Eric Granger , Yazid Attabi , Ghyslain Gagnon

Subjective room acoustics impressions play an important role for the performance and reception of music in concert venues and auralizations. Therefore, room acoustics since the 20th century dealt with the relationship between objective,…

Sound · Computer Science 2025-11-13 Leonie Böhlke , Tim Ziemer , Rolf Bader

This article introduces the Stochastic Texture Difference method for analyzing data at prescribed spatial and value scales. This method relies on constrained random walks around each pixel, describing how nearby image values typically…

Computer Vision and Pattern Recognition · Computer Science 2015-10-06 Nicolas Brodu , Hussein Yahia

Filters from the Gammatone family are often used to model auditory signal processing, but the filter constant values used to mimic human hearing are largely set to values based on historical psychoacoustic data collected several decades…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-13 Samiya A Alkhairy

The identification of siren sounds in urban soundscapes is a crucial safety aspect for smart vehicles and has been widely addressed by means of neural networks that ensure robustness to both the diversity of siren signals and the strong and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-16 Stefano Damiano , Thomas Dietzen , Toon van Waterschoot

In this paper, we examine the parameter estimation performance of three well-known sinusoidal models for speech and audio. The first one is the standard Sinusoidal Model (SM), which is based on the Fast Fourier Transform (FFT). The second…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-04 George P. Kafentzis

Systems for synthesizer sound matching, which automatically set the parameters of a synthesizer to emulate an input sound, have the potential to make the process of synthesizer programming faster and easier for novice and experienced…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-24 Fred Bruford , Frederik Blang , Shahan Nercessian

Discrimination mitigation within machine learning (ML) models could be complicated because multiple factors may be interwoven hierarchically and historically. Yet few existing fairness measures can capture the discrimination level within ML…

Machine Learning · Computer Science 2025-05-20 Yijun Bian , Yujie Luo , Ping Xu

Perceptual audio quality measurement systems algorithmically analyze the output of audio processing systems to estimate possible perceived quality degradation using perceptual models of human audition. In this manner, they save the time and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-14 Pablo M. Delgado , Jürgen Herre

Evaluation metrics in image synthesis play a key role to measure performances of generative models. However, most metrics mainly focus on image fidelity. Existing diversity metrics are derived by comparing distributions, and thus they…

Computer Vision and Pattern Recognition · Computer Science 2022-06-28 Jiyeon Han , Hwanil Choi , Yunjey Choi , Junho Kim , Jung-Woo Ha , Jaesik Choi