中文
相关论文

相关论文: Audio Impairment Recognition Using a Correlation-B…

200 篇论文

Long (> 200 ms) audio inpainting, to recover a long missing part in an audio segment, could be widely applied to audio editing tasks and transmission loss recovery. It is a very challenging problem due to the high dimensional, complex and…

声音 · 计算机科学 2019-11-18 Ya-Liang Chang , Kuan-Ying Lee , Po-Yu Wu , Hung-yi Lee , Winston Hsu

Audio-to-score alignment is an important pre-processing step for in-depth analysis of classical music. In this paper, we apply novel transposition-invariant audio features to this task. These low-dimensional features represent local pitch…

声音 · 计算机科学 2018-07-20 Andreas Arzt , Stefan Lattner

Identifying emotion from speech is a non-trivial task pertaining to the ambiguous definition of emotion itself. In this work, we adopt a feature-engineering based approach to tackle the task of speech emotion recognition. Formalizing our…

机器学习 · 计算机科学 2019-04-15 Gaurav Sahu

Audio deepfake detection is an emerging active topic. A growing number of literatures have aimed to study deepfake detection algorithms and achieved effective performance, the problem of which is far from being solved. Although there are…

声音 · 计算机科学 2023-08-30 Jiangyan Yi , Chenglong Wang , Jianhua Tao , Xiaohui Zhang , Chu Yuan Zhang , Yan Zhao

Audio fingerprinting is a well-established solution for song identification from short recording excerpts. Popular methods rely on the extraction of sparse representations, generally spectral peaks, and have proven to be accurate, fast, and…

声音 · 计算机科学 2023-10-31 Kamil Akesbi , Dorian Desblancs , Benjamin Martin

Higher criticism is a method for detecting signals that are both sparse and weak. Although first proposed in cases where the noise variables are independent, higher criticism also has reasonable performance in settings where those variables…

统计理论 · 数学 2010-10-05 Peter Hall , Jiashun Jin

Learning from Preferences in Reinforcement Learning (PbRL) has gained attention recently, as it serves as a natural fit for complicated tasks where the reward function is not easily available. However, preferences often come with…

机器学习 · 计算机科学 2026-03-19 Yuxuan Li , Harshith Reddy Kethireddy , Srijita Das

This paper addresses the problem of correlation estimation in sets of compressed images. We consider a framework where images are represented under the form of linear measurements due to low complexity sensing or security requirements. We…

计算机视觉与模式识别 · 计算机科学 2011-12-20 Vijayaraghavan Thirumalai , Pascal Frossard

Self-supervised representations excel at many vision and speech tasks, but their potential for audio-visual deepfake detection remains underexplored. Unlike prior work that uses these features in isolation or buried within complex…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Dragos-Alexandru Boldisor , Stefan Smeu , Dan Oneata , Elisabeta Oneata

This paper presents the Deep learning-based Perceptual Audio Quality metric (DeePAQ) for evaluating general audio quality. Our approach leverages metric learning together with the music foundation model MERT, guided by surrogate labels, to…

音频与语音处理 · 电气工程与系统科学 2025-10-15 Guanxin Jiang , Andreas Brendel , Pablo M. Delgado , Jürgen Herre

Audio fingerprinting systems must efficiently and robustly identify query snippets in an extensive database. To this end, state-of-the-art systems use deep learning to generate compact audio fingerprints. These systems deploy indexing…

音频与语音处理 · 电气工程与系统科学 2023-01-20 Anup Singh , Kris Demuynck , Vipul Arora

A key function of auditory cognition is the association of characteristic sounds with their corresponding semantics over time. Humans attempting to discriminate between fine-grained audio categories, often replay the same discriminative…

声音 · 计算机科学 2023-03-14 Alexandros Stergiou , Dima Damen

Understanding the internal mechanisms of large audio-language models (LALMs) is crucial for interpreting their behavior and improving performance. This work presents the first in-depth analysis of how LALMs internally perceive and recognize…

计算与语言 · 计算机科学 2025-08-26 Chih-Kai Yang , Neo Ho , Yi-Jyun Lee , Hung-yi Lee

We propose the Neuralogram -- a deep neural network based representation for understanding audio signals which, as the name suggests, transforms an audio signal to a dense, compact representation based upon embeddings learned via a neural…

声音 · 计算机科学 2019-04-11 Prateek Verma , Chris Chafe , Jonathan Berger

Fake audio detection is an emerging active topic. A growing number of literatures have aimed to detect fake utterance, which are mostly generated by Text-to-speech (TTS) or voice conversion (VC). However, countermeasures against…

声音 · 计算机科学 2024-09-02 Hao Gu , JiangYan Yi , Chenglong Wang , Yong Ren , Jianhua Tao , Xinrui Yan , Yujie Chen , Xiaohui Zhang

Machine learning techniques are an active area of research for speech enhancement for hearing aids, with one particular focus on improving the intelligibility of a noisy speech signal. Recent work has shown that feature encodings from…

声音 · 计算机科学 2024-07-19 Robert Sutherland , George Close , Thomas Hain , Stefan Goetze , Jon Barker

Accurately interpreting cardiac auscultation signals plays a crucial role in diagnosing and managing cardiovascular diseases. However, the paucity of labelled data inhibits classification models' training. Researchers have turned to…

声音 · 计算机科学 2025-06-18 Leigh Abbott , Milan Marocchi , Matthew Fynn , Yue Rong , Sven Nordholm

We propose a benchmark for evaluating compositionality in audio representations. Audio compositionality refers to representing sound scenes in terms of constituent sources and attributes, and combining them systematically. While central to…

声音 · 计算机科学 2026-03-17 Chuyang Chen , Bea Steers , Brian McFee , Juan Bello

We propose a novel method to use both audio and a low-resolution image to perform extreme face super-resolution (a 16x increase of the input size). When the resolution of the input image is very low (e.g., 8x8 pixels), the loss of…

计算机视觉与模式识别 · 计算机科学 2020-04-03 Givi Meishvili , Simon Jenni , Paolo Favaro

Research on speech processing has traditionally considered the task of designing hand-engineered acoustic features (feature engineering) as a separate distinct problem from the task of designing efficient machine learning (ML) models to…

声音 · 计算机科学 2021-09-27 Siddique Latif , Rajib Rana , Sara Khalifa , Raja Jurdak , Junaid Qadir , Björn W. Schuller