English
Related papers

Related papers: The iNaturalist Sounds Dataset

200 papers

This paper describes a data collection campaign and the resulting dataset derived from smartphone sensors characterizing the daily life activities of 3 volunteers in a period of two weeks. The dataset is released as a collection of CSV…

Human-Computer Interaction · Computer Science 2023-07-10 Mattia Giovanni Campana , Franca Delmastro

As speech generation technology advances, the risk of misuse through deepfake audio has become a pressing concern, which underscores the critical need for robust detection systems. However, many existing speech deepfake datasets are limited…

Sound · Computer Science 2025-07-30 Wen Huang , Yanmei Gu , Zhiming Wang , Huijia Zhu , Yanmin Qian

Sounds are essential to how humans perceive and interact with the world and are captured in recordings and shared on the Internet on a minute-by-minute basis. These recordings, which are predominantly videos, constitute the largest archive…

Sound · Computer Science 2023-03-31 Benjamin Elizalde , Rohan Badlani , Ankit Shah , Anurag Kumar , Bhiksha Raj

Generative modeling offers new opportunities for bioacoustics, enabling the synthesis of realistic animal vocalizations that could support biomonitoring efforts and supplement scarce data for endangered species. However, directly generating…

Sound · Computer Science 2025-09-03 Tianyu Song , Ton Viet Ta

Visual analysis of complex fish habitats is an important step towards sustainable fisheries for human consumption and environmental protection. Deep Learning methods have shown great promise for scene analysis when trained on large-scale…

Computer Vision and Pattern Recognition · Computer Science 2020-08-31 Alzayat Saleh , Issam H. Laradji , Dmitry A. Konovalov , Michael Bradley , David Vazquez , Marcus Sheaves

Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three-dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR,…

Certain environmental noises have been associated with negative developmental outcomes for infants and young children. Though classifying or tagging sound events in a domestic environment is an active research area, previous studies focused…

1. Natural sounds have been recorded for millions of hours over the previous decades using passive acoustic monitoring. Improvements in deep learning models have vastly accelerated the analysis of large portions of this data. While new…

Machine Learning · Computer Science 2026-04-14 Vincent S. Kather , Sylvain Haupert , Burooj Ghani , Dan Stowell

Animal vocalization denoising is a task similar to human speech enhancement, which is relatively well-studied. In contrast to the latter, it comprises a higher diversity of sound production mechanisms and recording environments, and this…

Being heavily reliant on animals, it is our ethical obligation to improve their well-being by understanding their needs. Several studies show that animal needs are often expressed through their faces. Though remarkable progress has been…

Computer Vision and Pattern Recognition · Computer Science 2019-09-12 Muhammad Haris Khan , John McDonagh , Salman Khan , Muhammad Shahabuddin , Aditya Arora , Fahad Shahbaz Khan , Ling Shao , Georgios Tzimiropoulos

The paper describes a dataset comprising indoor environmental factors such as temperature, humidity, air quality, and noise levels. The data was collected from 10 sensing devices installed in various locations within three single-family…

Signal Processing · Electrical Eng. & Systems 2023-10-09 Sheik Murad Hassan Anik , Xinghua Gao , Na Meng

This paper introduces VoxSim, a dataset of perceptual voice similarity ratings. Recent efforts to automate the assessment of speech synthesis technologies have primarily focused on predicting mean opinion score of naturalness, leaving…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-29 Junseok Ahn , Youkyum Kim , Yeunju Choi , Doyeop Kwak , Ji-Hoon Kim , Seongkyu Mun , Joon Son Chung

Environmental sounds like footsteps, keyboard typing, or dog barking carry rich information and emotional context, making them valuable for designing haptics in user applications. Existing audio-to-vibration methods, however, rely on…

Human-Computer Interaction · Computer Science 2026-01-27 Yinan Li , Hasti Seifi

Insects play such a crucial role in ecosystems that a shift in demography of just a few species can have devastating consequences at environmental, social and economic levels. Despite this, evaluation of insect demography is strongly…

Computer Vision and Pattern Recognition · Computer Science 2022-11-04 Léonard Boussioux , Tomás Giro-Larraz , Charles Guille-Escuret , Mehdi Cherti , Balázs Kégl

Large-scale in-the-wild speech datasets have become more prevalent in recent years due to increased interest in models that can learn useful features from unlabelled data for tasks such as speech recognition or synthesis. These datasets…

Obtaining large-scale human-labeled datasets to train acoustic representation models is a very challenging task. On the contrary, we can easily collect data with machine-generated labels. In this work, we propose to exploit…

Computer Vision and Pattern Recognition · Computer Science 2020-01-03 Shaoyong Jia , Xin Shu , Yang Yang , Dawei Liang , Qiyue Liu , Junhui Liu

This paper is an investigation into aspects of an audio classification pipeline that will be appropriate for the monitoring of bird species on edges devices. These aspects include transfer learning, data augmentation and model optimization.…

Sound · Computer Science 2021-08-11 David Behr , Ciira wa Maina , Vukosi Marivate

Our fight against false information is spearheaded by fact-checkers. They investigate the veracity of claims and document their findings as fact-checking reports. With the rapid increase in the amount of false information circulating…

Information Retrieval · Computer Science 2025-05-15 Enes Altuncu , Can Başkent , Sanjay Bhattacherjee , Shujun Li , Dwaipayan Roy

In this paper, ensembles of classifiers that exploit several data augmentation techniques and four signal representations for training Convolutional Neural Networks (CNNs) for audio classification are presented and tested on three freely…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-18 Loris Nanni , Gianluca Maguolo , Sheryl Brahnam , Michelangelo Paci

This paper introduces the acoustic scene classification task of DCASE 2018 Challenge and the TUT Urban Acoustic Scenes 2018 dataset provided for the task, and evaluates the performance of a baseline system in the task. As in previous years…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-12 Annamaria Mesaros , Toni Heittola , Tuomas Virtanen