English
Related papers

Related papers: Positive-Unlabelled Active Learning to Curate a Da…

200 papers

Certified machine unlearning aims to provably remove the influence of a deletion set $U$ from a model trained on a dataset $S$, by producing an unlearned output that is statistically indistinguishable from retraining on the retain set…

Machine Learning · Computer Science 2026-03-04 Carolin Heinzler , Kasra Malihi , Amartya Sanyal

A challenge in marine bioacoustic analysis is the detection of animal signals, like calls, whistles and clicks, for behavioral studies. Manual labeling is too time-consuming to process sufficient data to get reasonable results. Thus, an…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-23 Christopher Hauer

Training data cleaning is a new application for generative model-based speech restoration (SR). This paper introduces Miipher-2, an SR model designed for million-hour scale data, for training data cleaning for large-scale generative models…

On 21-22 November 2019, about 30 researchers gathered in Victoria, BC, Canada, for the workshop "Detection and Classification in Marine Bioacoustics with Deep Learning" organized by MERIDIAN and hosted by Ocean Networks Canada. The workshop…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-20 Fabio Frazao , Bruno Padovese , Oliver S. Kirsebom

Harnessing the power of Artificial Intelligence (AI) and m-health towards detecting new bio-markers indicative of the onset and progress of respiratory abnormalities/conditions has greatly attracted the scientific and research interest…

We present the Ultracool dwarf Science with MachIne LEarning (USMILE), a program developing machine-learning tools for the discovery and characterization of ultracool dwarfs. We introduce USMILE Avocado, a spectral classification framework…

Solar and Stellar Astrophysics · Physics 2025-10-21 Zhoujian Zhang , Yanxia Li

The sustainability of the ocean ecosystem is threatened by increased levels of sound pollution, making monitoring crucial to understand its variability and impact. Passive acoustic monitoring (PAM) systems collect a large amount of…

Sound · Computer Science 2025-05-27 Hilde I Hummel , Sandjai Bhulai , Burooj Ghani , Rob van der Mei

Machine Listening focuses on developing technologies to extract relevant information from audio signals. A critical aspect of these projects is the acquisition and labeling of contextualized data, which is inherently complex and requires…

Sound · Computer Science 2024-10-10 Javier Naranjo-Alcazar , Jordi Grau-Haro , Ruben Ribes-Serrano , Pedro Zuccarello

OpenAI Whisper is a family of robust Automatic Speech Recognition (ASR) models trained on 680,000 hours of audio. However, its encoder-decoder architecture, trained with a sequence-to-sequence objective, lacks native support for streaming…

Classification accuracy provided by a machine learning model depends a lot on the feature set used in the learning process. Feature Selection (FS) is an important and challenging pre-processing technique which helps to identify only the…

Machine Learning · Computer Science 2020-09-01 Ritam Guha , Manosij Ghosh , Shyok Mutsuddi , Ram Sarkar , Seyedali Mirjalili

Continual learning (CL) is crucial for language models to dynamically adapt to the evolving real-world demands. To mitigate the catastrophic forgetting problem in CL, data replay has been proven a simple and effective strategy, and the…

Computation and Language · Computer Science 2024-11-12 Jinghan He , Haiyun Guo , Kuan Zhu , Zihan Zhao , Ming Tang , Jinqiao Wang

The vast amounts of audio data collected in Sound Event Detection (SED) applications require efficient annotation strategies to enable supervised learning. Manual labeling is expensive and time-consuming, making Active Learning (AL) a…

Sound · Computer Science 2025-03-05 Richard Lindholm , Oscar Marklund , Olof Mogren , John Martinsson

In recent years, there has been considerable progress in research on human activity recognition using data from wearable sensors. This technology also has potential in the context of animal welfare in livestock science. In this paper, we…

Machine Learning · Computer Science 2024-08-26 Oshana Dissanayake , Lucile Riaboff , Sarah E. McPherson , Emer Kennedy , Pádraig Cunningham

Ocean renewable energy, particularly wave energy, has emerged as a pivotal component for diversifying the global energy portfolio, reducing dependence on fossil fuels, and mitigating climate change impacts. This study delves into the…

Neural and Evolutionary Computing · Computer Science 2023-09-20 Hossein Mehdipour , Erfan Amini , Seyed Taghi Naeeni , Mehdi Neshat

In the paper a method for automatic classification of signals received by EKB and MAGW ISTP SB RAS coherent scatter radars (8-20MHz operating frequency) during 2021 is described. The method is suitable for automatic physical interpretation…

We present the Geometric-Wave Acoustic (GWA) dataset, a large-scale audio dataset of about 2 million synthetic room impulse responses (IRs) and their corresponding detailed geometric and simulation configurations. Our dataset samples…

Sound · Computer Science 2022-06-22 Zhenyu Tang , Rohith Aralikatti , Anton Ratnarajah , Dinesh Manocha

We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoken language in Ethiopia and neighboring regions. The dataset was collected through a crowd-sourcing initiative, encompassing a diverse range…

Computation and Language · Computer Science 2025-02-04 Turi Abu , Ying Shi , Thomas Fang Zheng , Dong Wang

This paper proposes an active learning system for sound event detection (SED). It aims at maximizing the accuracy of a learned SED model with limited annotation effort. The proposed system analyzes an initially unlabeled audio dataset, from…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-10 Shuyang Zhao , Toni Heittola , Tuomas Virtanen

This work introduces the Cleanformer, a streaming multichannel neural based enhancement frontend for automatic speech recognition (ASR). This model has a conformer-based architecture which takes as inputs a single channel each of raw and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-05 Joseph Caroselli , Arun Narayanan , Nathan Howard , Tom O'Malley

Effective monitoring of whale populations is critical for conservation, but traditional survey methods are expensive and difficult to scale. While prior work has shown that whales can be identified in very high-resolution (VHR) satellite…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Caleb Robinson , Kimberly T. Goetz , Christin B. Khan , Meredith Sackett , Kathleen Leonard , Rahul Dodhia , Juan M. Lavista Ferres