中文
相关论文

相关论文: The Orchive : Data mining a massive bioacoustic ar…

200 篇论文

We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed…

声音 · 计算机科学 2022-07-18 Zhongweiyang Xu , Romit Roy Choudhury

A room's acoustic properties are a product of the room's geometry, the objects within the room, and their specific positions. A room's acoustic properties can be characterized by its impulse response (RIR) between a source and listener…

声音 · 计算机科学 2024-01-17 Mason Wang , Samuel Clarke , Jui-Hsien Wang , Ruohan Gao , Jiajun Wu

This work focuses on reliable detection and segmentation of bird vocalizations as recorded in the open field. Acoustic detection of avian sounds can be used for the automatized monitoring of multiple bird taxa and querying in long-term…

音频与语音处理 · 电气工程与系统科学 2017-11-20 Lefteris Fanioudakis , Ilyas Potamitis

We introduce EPIC-SOUNDS, a large-scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos. We propose an annotation pipeline where annotators temporally label…

声音 · 计算机科学 2025-07-17 Jaesung Huh , Jacob Chalk , Evangelos Kazakos , Dima Damen , Andrew Zisserman

Identification of bird species from audio records is one of the challenging tasks due to the existence of multiple species in the same recording, noise in the background, and long-term recording. Besides, choosing a proper acoustic feature…

声音 · 计算机科学 2022-01-04 Nahian Ibn Hasan

Spoken dialogue is a primary source of information in videos; therefore, accurately identifying who spoke what and when is essential for deep video understanding. We introduce D-ORCA, a \textbf{d}ialogue-centric \textbf{o}mni-modal large…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Changli Tang , Tianyi Wang , Fengyun Rao , Jing Lyu , Chao Zhang

We introduce a novel approach to studying animal behaviour and the context in which it occurs, through the use of microphone backpacks carried on the backs of individual free-flying birds. These sensors are increasingly used by animal…

声音 · 计算机科学 2016-12-19 Dan Stowell , Emmanouil Benetos , Lisa F. Gill

Research into the prediction and analysis of perceived audio quality is hampered by the scarcity of openly available datasets of audio signals accompanied by corresponding subjective quality scores. To address this problem, we present the…

音频与语音处理 · 电气工程与系统科学 2024-01-02 Matteo Torcoli , Chih-Wei Wu , Sascha Dick , Phillip A. Williams , Mhd Modar Halimeh , William Wolcott , Emanuel A. P. Habets

It is easier to hear birds than see them, however, they still play an essential role in nature and they are excellent indicators of deteriorating environmental quality and pollution. Recent advances in Machine Learning and Convolutional…

声音 · 计算机科学 2021-07-13 Marcos V. Conde , Kumar Shubham , Prateek Agnihotri , Nitin D. Movva , Szilard Bessenyei

In this work, we propose a supervised, convex representation based audio hashing framework for bird species classification. The proposed framework utilizes archetypal analysis, a matrix factorization technique, to obtain convex-sparse…

音频与语音处理 · 电气工程与系统科学 2019-02-08 Anshul Thakur , Pulkit Sharma , Vinayak Abrol , Padmanabhan Rajan

Our goal is to collect a large-scale audio-visual dataset with low label noise from videos in the wild using computer vision techniques. The resulting dataset can be used for training and evaluating audio recognition models. We make three…

计算机视觉与模式识别 · 计算机科学 2020-09-28 Honglie Chen , Weidi Xie , Andrea Vedaldi , Andrew Zisserman

We present a novel approach to automatically detect and classify great ape calls from continuous raw audio recordings collected during field research. Our method leverages deep pretrained and sequential neural networks, including wav2vec…

音频与语音处理 · 电气工程与系统科学 2024-06-24 Zifan Jiang , Adrian Soldati , Isaac Schamberg , Adriano R. Lameira , Steven Moran

We present a framework for detecting blue whale vocalisations from acoustic submarine recordings. The proposed methodology comprises three stages: i) a preprocessing step where the audio recordings are conditioned through normalisation,…

音频与语音处理 · 电气工程与系统科学 2021-10-06 Bryan Sagredo , Sonia Español-Jiménez , Felipe Tobar

On 21-22 November 2019, about 30 researchers gathered in Victoria, BC, Canada, for the workshop "Detection and Classification in Marine Bioacoustics with Deep Learning" organized by MERIDIAN and hosted by Ocean Networks Canada. The workshop…

音频与语音处理 · 电气工程与系统科学 2020-02-20 Fabio Frazao , Bruno Padovese , Oliver S. Kirsebom

The convergence of IoT sensing, edge computing, and machine learning is transforming precision livestock farming. Yet bioacoustic data streams remain underused because of computational complexity and ecological validity challenges. We…

声音 · 计算机科学 2025-10-17 Mayuri Kate , Suresh Neethirajan

Cardiac auscultation is one of the most cost-effective techniques used to detect and identify many heart conditions. Computer-assisted decision systems based on auscultation can support physicians in their decisions. Unfortunately, the…

Studying the vocalisations of wild animals can be a challenge due to the limitations of traditional computational methods, which often are time-consuming and lack reproducibility. Here, I present pykanto, a new software package that…

声音 · 计算机科学 2023-06-12 Nilo Merino Recalde

Speech recognition in highly-reverberant real environments remains a major challenge. An evaluation dataset for this task is needed. This report describes the generation of the Highly-Reverberant Real Environment database (HRRE). This…

音频与语音处理 · 电气工程与系统科学 2018-03-28 Juan Pablo Escudero , Victor Poblete , José Novoa , Jorge Wuth , Josué Fredes , Rodrigo Mahu , Richard Stern , Néstor Becerra Yoma

The Ricordi archive, a prestigious collection of significant musical manuscripts from renowned opera composers such as Donizetti, Verdi and Puccini, has been digitized. This process has allowed us to automatically extract samples that…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Federico Simonetta , Rishav Mondal , Luca Andrea Ludovico , Stavros Ntalampiras

This paper introduces the Voices Obscured In Complex Environmental Settings (VOICES) corpus, a freely available dataset under Creative Commons BY 4.0. This dataset will promote speech and signal processing research of speech recorded by…