English
Related papers

Related papers: Detection of manatee vocalisations using the Audio…

200 papers

Source separation is the process of isolating individual sounds in an auditory mixture of multiple sounds [1], and has a variety of applications ranging from speech enhancement and lyric transcription [2] to digital audio production for…

Sound · Computer Science 2024-12-10 Bradford Derby , Lucas Dunker , Samarth Galchar , Shashank Jarmale , Akash Setti

We introduce a novel approach to studying animal behaviour and the context in which it occurs, through the use of microphone backpacks carried on the backs of individual free-flying birds. These sensors are increasingly used by animal…

Sound · Computer Science 2016-12-19 Dan Stowell , Emmanouil Benetos , Lisa F. Gill

The North Atlantic right whale (Eubalaena glacialis) is an endangered species. These whales continuously suffer from deadly vessel impacts alongside the eastern coast of North America. There have been countless efforts to save the remaining…

Machine Learning · Computer Science 2013-06-10 Rami Abousleiman , Guangzhi Qu , Osamah Rawashdeh

Monoaural audio source separation is a challenging research area in machine learning. In this area, a mixture containing multiple audio sources is given, and a model is expected to disentangle the mixture into isolated atomic sources. In…

Machine Learning · Computer Science 2019-11-25 Amir Zadeh , Tianjun Ma , Soujanya Poria , Louis-Philippe Morency

Many species have evolved advanced non-visual perception while artificial systems fall behind. Radar and ultrasound complement camera-based vision but they are often too costly and complex to set up for very limited information gain. In…

Computer Vision and Pattern Recognition · Computer Science 2020-03-20 Jesper Haahr Christensen , Sascha Hornauer , Stella Yu

In the domain of Music Information Retrieval (MIR), Automatic Music Transcription (AMT) emerges as a central challenge, aiming to convert audio signals into symbolic notations like musical notes or sheet music. This systematic review…

Sound · Computer Science 2024-06-24 Fatemeh Jamshidi , Gary Pike , Amit Das , Richard Chapman

Transformers and State-Space Models (SSMs) have advanced audio classification by modeling spectrograms as sequences of patches. However, existing models such as the Audio Spectrogram Transformer (AST) and Audio Mamba (AuM) adopt square…

Sound · Computer Science 2025-09-01 Aditya Makineni , Baocheng Geng , Qing Tian

Understanding the internal representations of large language models is crucial for ensuring their reliability and safety, with sparse autoencoders (SAEs) emerging as a promising interpretability approach. However, current SAE training…

Machine Learning · Computer Science 2025-10-13 T. Ed Li , Junyu Ren

In this study, we evaluate the efficacy of the Mamba architecture bioacoustics by introducing BioMamba, a Mamba-based audio representation model for wildlife sounds. We pre-train a BioMamba using self-supervised learning on a large audio…

Sound · Computer Science 2026-04-21 Chengyu Tang , Sanjeev Baskiyar

Benthic habitat mapping is fundamental for understanding marine ecosystems, guiding conservation efforts, and supporting sustainable resource management. Yet, the scarcity of large, annotated datasets limits the development and benchmarking…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Hayat Rajani , Valerio Franchi , Borja Martinez-Clavel Valles , Raimon Ramos , Rafael Garcia , Nuno Gracias

Different machines can exhibit diverse frequency patterns in their emitted sound. This feature has been recently explored in anomaly sound detection and reached state-of-the-art performance. However, existing methods rely on the manual or…

Sound · Computer Science 2023-09-07 Hejing Zhang , Jian Guan , Qiaoxi Zhu , Feiyang Xiao , Youde Liu

Binaural Audio Telepresence (BAT) aims to encode the acoustic scene at the far end into binaural signals for the user at the near end. BAT encompasses an immense range of applications that can vary between two extreme modes of Immersive BAT…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-15 Yicheng Hsu , Mingsian R. Bai

Methods for cetacean research include photo-identification (photo-id) and passive acoustic monitoring (PAM) which generate thousands of images per expedition that are currently hand categorised by researchers into the individual dolphins…

Computer Vision and Pattern Recognition · Computer Science 2019-08-08 Cameron Trotter , Georgia Atkinson , Matthew Sharpe , A. Stephen McGough , Nick Wright , Per Berggren

Marine mammal communication is a complex field, hindered by the diversity of vocalizations and environmental factors. The Watkins Marine Mammal Sound Database (WMMD) constitutes a comprehensive labeled dataset employed in machine learning…

Signal Processing · Electrical Eng. & Systems 2024-06-27 Alessandro Licciardi , Davide Carbone

As deepfake speech becomes common and hard to detect, it is vital to trace its source. Recent work on audio deepfake source tracing (ST) aims to find the origins of synthetic or manipulated speech. However, ST models must adapt to learn new…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-21 Yang Xiao , Rohan Kumar Das

Autonomous recording units and passive acoustic monitoring present minimally intrusive methods of collecting bioacoustics data. Combining this data with species agnostic bird activity detection systems enables the monitoring of activity…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-04 Mark Anderson , Naomi Harte

Marine mammals are increasingly vulnerable to human disturbance and climate change. Their diving behavior leads to limited visual access during data collection, making studying the abundance and distribution of marine mammals challenging.…

End-to-end Speech-to-text Translation (E2E-ST), which directly translates source language speech to target language text, is widely useful in practice, but traditional cascaded approaches (ASR+MT) often suffer from error propagation in the…

Computation and Language · Computer Science 2021-02-10 Junkun Chen , Mingbo Ma , Renjie Zheng , Liang Huang

The use of auditory masking has long been of interest in psychoacoustics and for engineering purposes, in order to cover sounds that are disruptive to humans or to species whose habitats overlap with ours. In most cases, we seek to minimize…

Biological Physics · Physics 2025-10-17 Justin Faber , Alexandros C Alampounti , Marcos Georgiades , Joerg T Albert , Dolores Bozovic

The lack of annotated training data in bioacoustics hinders the use of large-scale neural network models trained in a supervised way. In order to leverage a large amount of unannotated audio data, we propose AVES (Animal Vocalization…

Sound · Computer Science 2022-10-27 Masato Hagiwara
‹ Prev 1 4 5 6 7 8 10 Next ›