English
Related papers

Related papers: BIRB: A Generalization Benchmark for Information R…

200 papers

General-purpose embedding is highly desirable for few-shot even zero-shot learning in many application scenarios, including audio tasks. In order to understand representations better, we conducted a thorough error analysis and visualization…

Sound · Computer Science 2023-03-08 Ankit Shah , Shuyi Chen , Kejun Zhou , Yue Chen , Bhiksha Raj

The performance of speaker diarization is strongly affected by its clustering algorithm at the test stage. However, it is known that clustering algorithms are sensitive to random noises and small variations, particularly when the clustering…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-25 Meng-Zhen Li , Xiao-Lei Zhang

Automatic detection and classification of animal sounds has many applications in biodiversity monitoring and animal behaviour. In the past twenty years, the volume of digitised wildlife sound available has massively increased, and automatic…

State-of-the-art machine learning methods exhibit limited compositional generalization. At the same time, there is a lack of realistic benchmarks that comprehensively measure this ability, which makes it challenging to find and evaluate…

What audio embedding approach generalizes best to a wide range of downstream tasks across a variety of everyday domains without fine-tuning? The aim of the HEAR benchmark is to develop a general-purpose audio representation that provides a…

Recent years has witnessed an increase in technologies that use speech for the sensing of the health of the talker. This survey paper proposes a general taxonomy of the technologies and a broad overview of current progress and challenges.…

Neurons and Cognition · Quantitative Biology 2024-08-12 Aki Härmä , Bert den Brinker , Ulf Grossekathofer , Okke Ouweltjes , Srikanth Nallanthighal , Sidharth Abrol , Vibhu Sharma

Animal vocalizations provide crucial insights for wildlife assessment, particularly in complex environments such as forests, aiding species identification and ecological monitoring. Recent advances in deep learning have enabled automatic…

Sound · Computer Science 2026-03-24 Risa Shinoda , Kaede Shiohara , Nakamasa Inoue , Hiroaki Santo , Fumio Okura

In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR)…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Xiulong Liu , Anurag Kumar , Paul Calamia , Sebastia V. Amengual , Calvin Murdock , Ishwarya Ananthabhotla , Philip Robinson , Eli Shlizerman , Vamsi Krishna Ithapu , Ruohan Gao

The estimation of reverberation time from real-world signals plays a central role in a wide range of applications. In many scenarios, acoustic conditions change over time which in turn requires the estimate to be updated continuously.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-11 Philipp Götz , Cagdas Tuna , Andreas Walther , Emanuël A. P. Habets

In ornithology, bird species are known to have variedit's widely acknowledged that bird species display diverse dialects in their calls across different regions. Consequently, computational methods to identify bird species onsolely through…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-14 Xin Jing , Luyang Zhang , Jiangjian Xie , Alexander Gebhard , Alice Baird , Bjoern Schuller

This paper is an investigation into aspects of an audio classification pipeline that will be appropriate for the monitoring of bird species on edges devices. These aspects include transfer learning, data augmentation and model optimization.…

Sound · Computer Science 2021-08-11 David Behr , Ciira wa Maina , Vukosi Marivate

Autoregressive (AR) modeling is invaluable in signal processing, in particular in speech and audio fields. Attempts in the literature can be found that regularize or constrain either the time-domain signal values or the AR coefficients,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-06 Ondřej Mokrý , Pavel Rajmic

Mastering fine-grained visual recognition, essential in many expert domains, can require that specialists undergo years of dedicated training. Modeling the progression of such expertize in humans remains challenging, and accurately…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Leonie Bossemeyer , Samuel Heinrich , Grant Van Horn , Oisin Mac Aodha

The Cram\'er-Rao bound (CRB), a well-known lower bound on the performance of any unbiased parameter estimator, has been used to study a wide variety of problems. However, to obtain the CRB, requires an analytical expression for the…

Machine Learning · Computer Science 2022-10-11 Hai Victor Habi , Hagit Messer , Yoram Bresler

Climate change is a major driver of biodiversity loss, changing the geographic range and abundance of many species. However, there remain significant knowledge gaps about the distribution of species, due principally to the amount of effort…

Computer Vision and Pattern Recognition · Computer Science 2023-05-03 Mélisande Teng , Amna Elmustafa , Benjamin Akera , Hugo Larochelle , David Rolnick

Detecting the presence of animal vocalisations in nature is essential to study animal populations and their behaviors. A recent development in the field is the introduction of the task known as few-shot bioacoustic sound event detection,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-28 Jinhua Liang , Ines Nolasco , Burooj Ghani , Huy Phan , Emmanouil Benetos , Dan Stowell

Speech emotion recognition (SER) has made significant strides with the advent of powerful self-supervised learning (SSL) models. However, the generalization of these models to diverse languages and emotional expressions remains a challenge.…

Computation and Language · Computer Science 2024-08-16 Mohamed Osman , Daniel Z. Kaplan , Tamer Nadeem

Most work in audio enhancement targets human speech, while bioacoustics is less studied due to noisy recordings and the distinct traits of animal sounds. To fill this gap, we adapt speech enhancement methods and build BioSEN, a model made…

Sound · Computer Science 2026-05-15 Tianyu Song , Ton Viet Ta , Ngamta Thamwattana , Hisako Nomura , Linh Thi Hoai Nguyen

Motivation: Public and private repositories of experimental data are growing to sizes that require dedicated methods for finding relevant data. To improve on the state of the art of keyword searches from annotations, methods for…

Machine Learning · Statistics 2016-01-08 Paul Blomstedt , Ritabrata Dutta , Sohan Seth , Alvis Brazma , Samuel Kaski

In the context of building acoustics and the acoustic diagnosis of an existing room, this paper introduces and investigates a new approach to estimate mean absorption coefficients solely from a room impulse response (RIR). This inverse…

Neural and Evolutionary Computing · Computer Science 2021-09-02 Cédric Foy , Antoine Deleforge , Diego Di Carlo
‹ Prev 1 4 5 6 7 8 10 Next ›