English
Related papers

Related papers: BioME: A Resource-Efficient Bioacoustic Foundation…

200 papers

Although deep learning has advanced automated electrocardiogram (ECG) diagnosis, prevalent supervised methods typically treat recordings as undifferentiated one-dimensional (1D) signals or two-dimensional (2D) images. This formulation…

Machine Learning · Computer Science 2026-01-13 Runze Ma , Caizhi Liao

In the field of acoustic scene analysis, this paper presents a novel approach to find spatio-temporal latent representations from in-the-wild audio data. By using WE-LIVE, an in-house collected dataset that includes audio recordings in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-11 Claudia Montero-Ramírez , Esther Rituerto-González , Carmen Peláez-Moreno

In response to the burgeoning global demand for seafood and the challenges of managing fish farms, we introduce an innovative IoT based environmental control system that integrates sensor technology and advanced machine learning decision…

Signal Processing · Electrical Eng. & Systems 2023-11-09 D. Dhinakaran , S. Gopalakrishnan , M. D. Manigandan , T. P. Anish

This paper demonstrates the potential of vibration-based Foundation Models (FMs), pre-trained with unlabeled sensing data, to improve the robustness of run-time inference in (a class of) IoT applications. A case study is presented featuring…

With the emergence of Artificial Intelligence (AI), new attention has been given to implement AI algorithms on resource constrained tiny devices to expand the application domain of IoT. Multimodal Learning has recently become very popular…

Machine Learning · Computer Science 2022-04-20 Hasib-Al Rashid , Pretom Roy Ovi , Carl Busart , Aryya Gangopadhyay , Tinoosh Mohsenin

Discovering structure in biological signals without supervision is a fundamental problem in computational intelligence, yet existing bioacoustic methods assume vocal production models or predefined semantic units, leaving non-vocal species…

Sound · Computer Science 2026-05-11 Hamze Hammami , Nidhal Abdulaziz

Passive acoustic monitoring (PAM) studies generate thousands of hours of audio, which may be used to monitor specific animal populations, conduct broad biodiversity surveys, detect threats such as poachers, and more. Machine learning…

Quantitative Methods · Quantitative Biology 2024-02-26 Amanda K. Navine , Tom Denton , Matthew J. Weldy , Patrick J. Hart

1. Many ecological decisions are slowed by the gap between collecting and analysing biodiversity data. Edge computing moves processing closer to the sensor, with edge artificial intelligence (AI) enabling on-device inference, reducing…

Computers and Society · Computer Science 2026-02-17 Aude Vuilliomenet , Kate E. Jones , Duncan Wilson

Discrete audio tokens derived from self-supervised learning models have gained widespread usage in speech generation. However, current practice of directly utilizing audio tokens poses challenges for sequence modeling due to the length of…

Sound · Computer Science 2024-01-17 Feiyu Shen , Yiwei Guo , Chenpeng Du , Xie Chen , Kai Yu

Interactions with virtual assistants typically start with a trigger phrase followed by a command. In this work, we explore the possibility of making these interactions more natural by eliminating the need for a trigger phrase. Our goal is…

Monitoring of bird populations has played a vital role in conservation efforts and in understanding biodiversity loss. The automation of this process has been facilitated by both sensing technologies, such as passive acoustic monitoring,…

Machine Learning · Computer Science 2021-08-23 Irina Tolkova , Brian Chu , Marcel Hedman , Stefan Kahl , Holger Klinck

Audio self-supervised learning (SSL) pre-training, which aims to learn good representations from unlabeled audio, has made remarkable progress. However, the extensive computational demands during pre-training pose a significant barrier to…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Wenxi Chen , Yuzhe Liang , Ziyang Ma , Zhisheng Zheng , Xie Chen

Universities are microcosms of urban ecosystems, with concentrated consumption patterns in food, transport, energy, and product usage. These environments not only contribute substantially to sustainability pressures but also provide a…

Human-Computer Interaction · Computer Science 2026-04-20 Caleb Adu , Neil Kapadia , Binhe Liu , Jonathan Randall , Sruthi Viswanathan

This technical report describes our submission to the ICME 2025 audio encoder challenge. Our submitted system is built on BEATs, a masked speech token prediction based audio encoder. We extend the BEATs model using 74,000 hours of data…

Human-robot interaction relies on a noise-robust audio processing module capable of estimating target speech from audio recordings impacted by environmental noise, as well as self-induced noise, so-called ego-noise. While external ambient…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-28 Huajian Fang , Niklas Wittmer , Johannes Twiefel , Stefan Wermter , Timo Gerkmann

Deep learning (DL) has greatly advanced audio classification, yet the field is limited by the scarcity of large-scale benchmark datasets that have propelled progress in other domains. While AudioSet is a pivotal step to bridge this gap as a…

Masked Autoencoders (MAEs) learn rich semantic representations in audio classification through an efficient self-supervised reconstruction task. However, general-purpose models fail to generalize well when applied directly to fine-grained…

Machine Learning · Computer Science 2025-08-20 Lukas Rauch , René Heinrich , Ilyass Moummad , Alexis Joly , Bernhard Sick , Christoph Scholz

Leveraging multimodal information from biosignals is vital for building a comprehensive representation of people's physical and mental states. However, multimodal biosignals often exhibit substantial distributional shifts between…

Machine Learning · Computer Science 2024-04-22 Ran Liu , Ellen L. Zippi , Hadi Pouransari , Chris Sandino , Jingping Nie , Hanlin Goh , Erdrin Azemi , Ali Moin

Neural network models for audio tasks, such as automatic speech recognition (ASR) and acoustic scene classification (ASC), are susceptible to noise contamination for real-life applications. To improve audio quality, an enhancement module,…

1. Automated analysis of bioacoustic recordings using machine learning (ML) methods has the potential to greatly scale biodiversity monitoring efforts. The use of ML for high-stakes applications, such as conservation research, demands a…