English
Related papers

Related papers: Full-Stack Bioacoustics: Field Kit to AI to Action…

200 papers

Audio recorded in real-world environments often contains a mixture of foreground speech and background environmental sounds. With rapid advances in text-to-speech, voice conversion, and other generation models, either component can now be…

Sound · Computer Science 2026-02-06 Xueping Zhang , Han Yin , Yang Xiao , Lin Zhang , Ting Dang , Rohan Kumar Das , Ming Li

In this white paper, we synthesize key points made during presentations and discussions from the AI-Assisted Decision Making for Conservation workshop, hosted by the Center for Research on Computation and Society at Harvard University on…

1. Passive acoustic monitoring of biodiversity is growing fast, as it offers an alternative to traditional aural point count surveys, with the possibility to deploy long-term acoustic surveys in large and complex natural environments.…

Physics and Society · Physics 2022-11-30 Sylvain Haupert , Frédéric Sèbe , Jérôme Sueur

While AI presents significant potential for enhancing music mixing and mastering workflows, current research predominantly emphasizes end-to-end automation or generation, often overlooking the collaborative and instructional dimensions…

Sound · Computer Science 2025-07-10 Michael Clemens , Ana Marasović

Automatic identification of animal species by their vocalization is an important and challenging task. Although many kinds of audio monitoring system have been proposed in the literature, they suffer from several disadvantages such as…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-25 Weitao Xu , Xiang Zhang , Lina Yao , Wanli Xue , Bo Wei

Analysis tools used in research laboratories, for sound synthesis, by musicians or sound engineers can be rather different. Discussion of the assumptions and of the limitations of these tools permits to propose a first tool as relevant and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-08 Laurent Millot

In the speech signal, acoustic landmarks identify times when the acoustic manifestations of the linguistically motivated distinctive features are most salient. Acoustic landmarks have been widely applied in various domains, including speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-26 Xiangyu Zhang , Daijiao Liu , Tianyi Xiao , Cihan Xiao , Tuende Szalay , Mostafa Shahin , Beena Ahmed , Julien Epps

Recently, sound recognition has been used to identify sounds, such as car and river. However, sounds have nuances that may be better described by adjective-noun pairs such as slow car, and verb-noun pairs such as flying insects, which are…

Sound · Computer Science 2018-01-10 Sebastian Sager , Benjamin Elizalde , Damian Borth , Christian Schulze , Bhiksha Raj , Ian Lane

BickGraphing is a browser based research tool that enables visual inspection of acoustic recordings. The tool was built in support of visualizing crop feeding pest sounds in support of the Insect Eavesdropper project; however, it is widely…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-27 Kayley Seow , Alexander Arovas , Grace Steinmetz , Emily Bick

Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given a natural language query, which provides a natural and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-03 Xubo Liu , Qiuqiang Kong , Yan Zhao , Haohe Liu , Yi Yuan , Yuzhuo Liu , Rui Xia , Yuxuan Wang , Mark D. Plumbley , Wenwu Wang

Identification and localization of sounds are both integral parts of computational auditory scene analysis. Although each can be solved separately, the goal of forming coherent auditory objects and achieving a comprehensive spatial scene…

Sound · Computer Science 2019-12-24 Ivo Trowitzsch , Christopher Schymura , Dorothea Kolossa , Klaus Obermayer

This work introduces a robotic dummy head that fuses the acoustic realism of conventional audiological mannequins with the mobility of robots. The proposed device is capable of moving, talking, and listening as people do, and can be used to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-08 Austin Lu , Kanad Sarkar , Yongjie Zhuang , Leo Lin , Ryan M Corey , Andrew C Singer

Insects represent half of all global biodiversity, yet many of the world's insects are disappearing, with severe implications for ecosystems and agriculture. Despite this crisis, data on insect diversity and abundance remain woefully…

Environmental sound scene and sound event recognition is important for the recognition of suspicious events in indoor and outdoor environments (such as nurseries, smart homes, nursing homes, etc.) and is a fundamental task involved in many…

Sound · Computer Science 2023-08-31 Nan Che , Chenrui Liu , Fei Yu

We demonstrate tools and applications developed based on the method of "sound safeguarding," which enables any sound to be used for acoustic measurements. We developed tools for preparation, interactive and real-time measurement, and report…

Sound · Computer Science 2025-07-29 Hideki Kawahara , Kohei Yatabe , Ken-Ichi Sakakibara

Spoofed audio, i.e. audio that is manipulated or AI-generated deepfake audio, is difficult to detect when only using acoustic features. Some recent innovative work involving AI-spoofed audio detection models augmented with phonetic and…

Sound · Computer Science 2024-10-22 Zahra Khanjani , Christine Mallinson , James Foulds , Vandana P Janeja

Obtaining data to train robust artificial intelligence (AI)-based models for species classification can be challenging, particularly for rare species. Data augmentation can boost classification accuracy by increasing the diversity of…

Sound · Computer Science 2025-12-16 Anthony Gibbons , Emma King , Ian Donohue , Andrew Parnell

This study investigates the potential of automated deep learning to enhance the accuracy and efficiency of multi-class classification of bird vocalizations, compared against traditional manually-designed deep learning models. Using the…

Machine Learning · Computer Science 2023-12-27 Giulio Tosato , Abdelrahman Shehata , Joshua Janssen , Kees Kamp , Pramatya Jati , Dan Stowell

Several animal species (e.g., bats, dolphins, and whales) and even visually impaired humans have the remarkable ability to perform echolocation: a biological sonar used to perceive spatial layout and locate objects in the world. We explore…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Ruohan Gao , Changan Chen , Ziad Al-Halah , Carl Schissler , Kristen Grauman

It is argued based on the results of both numerical modelling and the experiments performed on an artificial substitute of a meadow that the sound emitted by animals living in a dense surrounding such as a meadow or shrubs can be used as a…