English
Related papers

Related papers: Positive-Unlabelled Active Learning to Curate a Da…

200 papers

Ecologists often use a hidden Markov model to decode a latent process, such as a sequence of an animal's behaviours, from an observed biologging time series. Modern technological devices such as video recorders and drones now allow…

Generative modeling offers new opportunities for bioacoustics, enabling the synthesis of realistic animal vocalizations that could support biomonitoring efforts and supplement scarce data for endangered species. However, directly generating…

Sound · Computer Science 2025-09-03 Tianyu Song , Ton Viet Ta

This report introduces Dolphin, a large-scale multilingual automatic speech recognition (ASR) model that extends the Whisper architecture to support a wider range of languages. Our approach integrates in-house proprietary and open-source…

Computation and Language · Computer Science 2025-03-27 Yangyang Meng , Jinpeng Li , Guodong Lin , Yu Pu , Guanbo Wang , Hu Du , Zhiming Shao , Yukai Huang , Ke Li , Wei-Qiang Zhang

Automated plankton image recognition is increasingly used in aquatic ecosystem monitoring, but deployed classifiers inevitably encounter unseen taxa and non-target particles. Open-set recognition methods are usually evaluated with…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Xi Chen , Eryuan Huang , Yingjun Xiao , Gang Fang

Passive acoustic monitoring (PAM) enables large-scale biodiversity assessment, but continuous recording generates large amounts of non-informative audio, creating challenges for storage, power consumption, and long-term edge deployment.…

Sound · Computer Science 2026-05-21 Muhammad Mun'im Ahmad Zabidi , Mohd Yamani Idna Idris , Norisma Idris

Sperm whales (Physeter macrocephalus) navigate underwater with a series of impulsive, click-like sounds known as echolocation clicks. These clicks are characterized by a multipulse structure (MPS) that serves as a distinctive pattern. In…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-03 Guy Gubnitsky , Roee Diamant

Effective monitoring of wildlife is critical for assessing biodiversity and ecosystem health, as declines in key species often signal significant environmental changes. Birds, particularly ground-nesting species, serve as important…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Carl Chalmers , Paul Fergus , Serge Wich , Steven N Longmore , Naomi Davies Walsh , Lee Oliver , James Warrington , Julieanne Quinlan , Katie Appleby

Birds are vital parts of ecosystems across the world and are an excellent measure of the quality of life on earth. Many bird species are endangered while others are already extinct. Ecological efforts in understanding and monitoring bird…

Multimedia · Computer Science 2022-11-16 Chandra Kanth Nagesh , Abhishek Purushothama

Continual Learning (CL) involves fine-tuning pre-trained models with new data while maintaining the performance on the pre-trained data. This is particularly relevant for expanding multilingual ASR (MASR) capabilities. However, existing CL…

Computation and Language · Computer Science 2024-09-30 Chin Yuen Kwok , Jia Qi Yip , Eng Siong Chng

Today, data collection has improved in various areas, and the medical domain is no exception. Auscultation, as an important diagnostic technique for physicians, due to the progress and availability of digital stethoscopes, lends itself well…

Wake word (WW) spotting is challenging in far-field not only because of the interference in signal transmission but also the complexity in acoustic environments. Traditional WW model training requires large amount of in-domain WW-specific…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-15 Yixin Gao , Yuriy Mishchenko , Anish Shah , Spyros Matsoukas , Shiv Vitaladevuni

We present ORACLE, the first hierarchical deep-learning model for real-time, context-aware classification of transient and variable astrophysical phenomena. ORACLE is a recurrent neural network with Gated Recurrent Units (GRUs), and has…

Instrumentation and Methods for Astrophysics · Physics 2025-12-04 Ved G. Shah , Alex Gagliano , Konstantin Malanchev , Gautham Narayan , Alex I. Malz , The LSST Dark Energy Science Collaboration

This paper introduces a novel framework for open-set speaker identification in household environments, playing a crucial role in facilitating seamless human-computer interactions. Addressing the limitations of current speaker models and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-25 Zhiyong Chen , Zhiqi Ai , Xinnuo Li , Shugong Xu

Light curves serve as a valuable source of information on stellar formation and evolution. With the rapid advancement of machine learning techniques, it can be effectively processed to extract astronomical patterns and information. In this…

Instrumentation and Methods for Astrophysics · Physics 2025-03-18 Yu-Yang Li , Yu Bai , Cunshi Wang , Mengwei Qu , Ziteng Lu , Roberto Soria , Jifeng Liu

Recent research has focused on enhancing the capability of smaller models through imitation learning, drawing on the outputs generated by large foundation models (LFMs). A number of issues impact the quality of these models, ranging from…

Computation and Language · Computer Science 2023-06-06 Subhabrata Mukherjee , Arindam Mitra , Ganesh Jawahar , Sahaj Agarwal , Hamid Palangi , Ahmed Awadallah

Automated detection of acoustic signals is crucial for effective monitoring of sound-producing animals and their habitats across ecologically relevant spatial and temporal scales. Recent advances in deep learning have made these approaches…

We compare self-supervised representation learning algorithms which either explicitly quantize the audio data or learn representations without quantization. We find the former to be more accurate since it builds a good vocabulary of the…

Computation and Language · Computer Science 2020-05-20 Alexei Baevski , Michael Auli , Abdelrahman Mohamed

Narwhal is one of the most mysterious marine mammals, due to its isolated habitat in the Arctic region. Tagging is a technology that has the potential to explore the activities of this species, where behavioral information can be collected…

For conversational large-vocabulary continuous speech recognition (LVCSR) tasks, up to about two thousand hours of audio is commonly used to train state of the art models. Collection of labeled conversational audio however, is prohibitively…

Computation and Language · Computer Science 2017-05-30 Shane Walker , Morten Pedersen , Iroro Orife , Jason Flaks

Models based on the Transformer architecture have seen widespread application across fields such as natural language processing, computer vision, and robotics, with large language models like ChatGPT revolutionizing machine understanding of…

Robotics · Computer Science 2024-07-24 Yujian Dong , Tianyu Wu , Chaoyang Song
‹ Prev 1 3 4 5 6 7 10 Next ›