English
Related papers

Related papers: Parameter Sensitivity of Deep-Feature based Evalua…

200 papers

The expressive nature of the voice provides a powerful medium for communicating sonic ideas, motivating recent research on methods for query by vocalisation. Meanwhile, deep learning methods have demonstrated state-of-the-art results for…

Multimedia · Computer Science 2018-02-15 Adib Mehrabi , Keunwoo Choi , Simon Dixon , Mark Sandler

Diffusion models have demonstrated remarkable success in generative tasks, including audio super-resolution (SR). In many applications like movie post-production and album mastering, substantial computational budgets are available for…

Sound · Computer Science 2025-08-05 Yizhu Jin , Zhen Ye , Zeyue Tian , Haohe Liu , Qiuqiang Kong , Yike Guo , Wei Xue

Neural audio codecs have gained recent popularity for their use in generative modeling as they offer high-fidelity audio reconstruction at low bitrates. While human listening studies remain the gold standard for assessing perceptual…

Sound · Computer Science 2025-11-26 Luca A. Lanzendörfer , Florian Grötschla

Fr\'echet Audio Distance (FAD) is the de facto standard for evaluating text-to-audio generation, yet its scores depend on the underlying encoder's embedding space. An encoder's training task dictates which acoustic features are preserved or…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-02 Wonwoo Jeong

Super-resolution results are usually measured by full-reference image quality metrics or human rating scores. However, these evaluation methods are general image quality measurement, and do not account for the nature of the super-resolution…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Sheng Cheng

Many speech processing methods based on deep learning require an automatic and differentiable audio metric for the loss function. The DPAM approach of Manocha et al. learns a full-reference metric trained directly on human judgments, and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-11 Pranay Manocha , Zeyu Jin , Richard Zhang , Adam Finkelstein

Psychoacoustical so-called "timbre spaces" map perceptual similarity ratings of instrument sounds onto low-dimensional embeddings via multidimensional scaling, but suffer from scalability issues and are incapable of generalization. Recent…

Sound · Computer Science 2025-07-11 Haokun Tian , Stefan Lattner , Charalampos Saitis

Based on the normal mode and ray theory, this article discusses the characteristics of surface sound source and reception at the surface layer, and explores depth estimation methods based on normal modes and rays, and proposes a depth…

This paper proposes attentive statistics pooling for deep speaker embedding in text-independent speaker verification. In conventional speaker embedding, frame-level features are averaged over all the frames of a single utterance to form an…

Audio and Speech Processing · Electrical Eng. & Systems 2019-02-27 Koji Okabe , Takafumi Koshinaka , Koichi Shinoda

In the development of spatial audio technologies, reliable and shared methods for evaluating audio quality are essential. Listening tests are currently the standard but remain costly in terms of time and resources. Several models predicting…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Adrien Llave , Emma Granier , Grégory Pallone

This article presents a review of typical techniques used in three distinct aspects of deep learning model development for audio generation. In the first part of the article, we provide an explanation of audio representations, beginning…

Sound · Computer Science 2024-06-04 Matej Božić , Marko Horvat

This study focuses on the perception of music performances when contextual factors, such as room acoustics and instrument, change. We propose to distinguish the concept of "performance" from the one of "interpretation", which expresses the…

Sound · Computer Science 2022-03-08 Federico Simonetta , Federico Avanzini , Stavros Ntalampiras

The evaluation of synthetic micro-structure images is an emerging problem as machine learning and materials science research have evolved together. Typical state of the art methods in evaluating synthetic images from generative models have…

Materials Science · Physics 2022-11-18 Devesh Shah , Anirudh Suresh , Alemayehu Admasu , Devesh Upadhyay , Kalyanmoy Deb

Accurately interpreting cardiac auscultation signals plays a crucial role in diagnosing and managing cardiovascular diseases. However, the paucity of labelled data inhibits classification models' training. Researchers have turned to…

Sound · Computer Science 2025-06-18 Leigh Abbott , Milan Marocchi , Matthew Fynn , Yue Rong , Sven Nordholm

One of the biggest challenges in multi-microphone applications is the estimation of the parameters of the signal model such as the power spectral densities (PSDs) of the sources, the early (relative) acoustic transfer functions of the…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-16 Andreas I. Koutrouvelis , Richard C. Hendriks , Richard Heusdens , Jesper Jensen

Multiple metrics have been introduced to measure fairness in various natural language processing tasks. These metrics can be roughly categorized into two categories: 1) \emph{extrinsic metrics} for evaluating fairness in downstream…

Computation and Language · Computer Science 2022-03-29 Yang Trista Cao , Yada Pruksachatkun , Kai-Wei Chang , Rahul Gupta , Varun Kumar , Jwala Dhamala , Aram Galstyan

While audio quality is a key performance metric for various audio processing tasks, including generative modeling, its objective measurement remains a challenge. Audio-Language Models (ALMs) are pre-trained on audio-text pairs that may…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-02 Soham Deshmukh , Dareen Alharthi , Benjamin Elizalde , Hannes Gamper , Mahmoud Al Ismail , Rita Singh , Bhiksha Raj , Huaming Wang

This paper proposes new acoustic feature signatures based on the multiscale fractal dimension (MFD), which are robust against the diversity of environmental sounds, for the content-based similarity search. The diversity of sound sources and…

Sound · Computer Science 2021-10-13 Motohiro Sunouchi , Masaharu Yoshioka

In this paper, we introduce a non-parametric texture similarity measure based on the singular value decomposition of the curvelet coefficients followed by a content-based truncation of the singular values. This measure focuses on images…

Image and Video Processing · Electrical Eng. & Systems 2019-02-01 Motaz Alfarraj , Yazeed Alaudah , Ghassan AlRegib

We present the Perceptimatic English Benchmark, an open experimental benchmark for evaluating quantitative models of speech perception in English. The benchmark consists of ABX stimuli along with the responses of 91 American…

Computation and Language · Computer Science 2020-05-08 Juliette Millet , Ewan Dunbar