中文
相关论文

相关论文: ASGIR: Audio Spectrogram Transformer Guided Classi…

200 篇论文

Saving rainforests is a key to halting adverse climate changes. In this paper, we introduce an innovative solution built on acoustic surveillance and machine learning technologies to help rainforest conservation. In particular, We propose…

声音 · 计算机科学 2019-08-22 Yuan Liu , Zhongwei Cheng , Jie Liu , Bourhan Yassin , Zhe Nan , Jiebo Luo

One solution to automatic speech recognition (ASR) of overlapping speakers is to separate speech and then perform ASR on the separated signals. Commonly, the separator produces artefacts which often degrade ASR performance. Addressing this…

Machine learning algorithms, when trained on audio recordings from a limited set of devices, may not generalize well to samples recorded using other devices with different frequency responses. In this work, a relatively straightforward…

声音 · 计算机科学 2021-05-26 Michał Kośmider

In this work, we consider applying machine learning to the analysis and compression of audio signals in the context of monitoring elephants in sub-Saharan Africa. Earth's biodiversity is increasingly under threat by sources of anthropogenic…

Audio carries richer information than text, including emotion, speaker traits, and environmental context, while also enabling lower-latency processing compared to speech-to-text pipelines. However, recent multimodal information retrieval…

声音 · 计算机科学 2026-04-23 Tong Zhao , Chenghao Zhang , Yutao Zhu , Zhicheng Dou

The identification of siren sounds in urban soundscapes is a crucial safety aspect for smart vehicles and has been widely addressed by means of neural networks that ensure robustness to both the diversity of siren signals and the strong and…

音频与语音处理 · 电气工程与系统科学 2024-09-16 Stefano Damiano , Thomas Dietzen , Toon van Waterschoot

Environmental Sound Classification (ESC) is an important and challenging problem, and feature representation is a critical and even decisive factor in ESC. Feature representation ability directly affects the accuracy of sound…

声音 · 计算机科学 2019-08-19 Tianhao Qiao , Shunqing Zhang , Zhichao Zhang , Shan Cao , Shugong Xu

Many ecosystems can undergo important qualitative changes, including sudden transitions to alternative stable states, in response to perturbations or increments in conditions. Such 'tipping points' are often preceded by declines in aspects…

种群与进化 · 定量生物学 2025-09-04 Neel P. Le Penru , Thomas M. Bury , Sarab S. Sethi , Robert M. Ewers , Lorenzo Picinali

Autonomous recording units and passive acoustic monitoring present minimally intrusive methods of collecting bioacoustics data. Combining this data with species agnostic bird activity detection systems enables the monitoring of activity…

音频与语音处理 · 电气工程与系统科学 2022-10-04 Mark Anderson , Naomi Harte

Speaker Identification refers to the process of identifying a person using one's voice from a collection of known speakers. Environmental noise, reverberation and distortion make the task of automatic speaker identification challenging as…

音频与语音处理 · 电气工程与系统科学 2025-09-01 Sabbir Ahmed , Nursadul Mamun , Md Azad Hossain

Birdsong often contains large amounts of rapid frequency modulation (FM). It is believed that the use or otherwise of FM is adaptive to the acoustic environment, and also that there are specific social uses of FM such as trills in…

声音 · 计算机科学 2015-09-22 Dan Stowell , Mark D. Plumbley

Voice activity detection is an essential pre-processing component for speech-related tasks such as automatic speech recognition (ASR). Traditional supervised VAD systems obtain frame-level labels from an ASR pipeline by using, e.g., a…

声音 · 计算机科学 2021-05-11 Heinrich Dinkel , Shuai Wang , Xuenan Xu , Mengyue Wu , Kai Yu

Effective conservation of maritime environments and wildlife management of endangered species require the implementation of efficient, accurate and scalable solutions for environmental monitoring. Ecoacoustics offers the advantages of…

声音 · 计算机科学 2025-07-29 Burla Nur Korkmaz , Roee Diamant , Gil Danino , Alberto Testolin

Studying the vocalisations of wild animals can be a challenge due to the limitations of traditional computational methods, which often are time-consuming and lack reproducibility. Here, I present pykanto, a new software package that…

声音 · 计算机科学 2023-06-12 Nilo Merino Recalde

Time-frequency representations of audio signals often resemble texture images. This paper derives a simple audio classification algorithm based on treating sound spectrograms as texture images. The algorithm is inspired by an earlier visual…

计算机视觉与模式识别 · 计算机科学 2008-09-29 Guoshen Yu , Jean-Jacques Slotine

Distortion products are tones produced through nonlinear effects of a system simultaneously detecting two or more frequencies. These combination tones are ubiquitous to vertebrate auditory systems and are generally regarded as byproducts of…

Language Identification (LID) systems are used to classify the spoken language from a given audio sample and are typically the first step for many spoken language processing tasks, such as Automatic Speech Recognition (ASR) systems. Without…

计算机视觉与模式识别 · 计算机科学 2017-08-17 Christian Bartz , Tom Herold , Haojin Yang , Christoph Meinel

Passive Acoustic Monitoring (PAM) is an efficient and non-invasive method for surveying ecosystems at a reduced cost. Typically, autonomous recorders allow the acquisition of vast bioacoustic datasets which are then analyzed. However, power…

声音 · 计算机科学 2026-05-06 Louis Lerbourg , Paul Peyret , Juliette Linossier , Marielle Malfante

Global biodiversity is declining at an unprecedented rate, yet little information is known about most species and how their populations are changing. Indeed, some 90% of Earth's species are estimated to be completely unknown. Machine…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Yuyan Chen , Nico Lang , B. Christian Schmidt , Aditya Jain , Yves Basset , Sara Beery , Maxim Larrivée , David Rolnick

The problem of pitch tracking has been extensively studied in the speech research community. The goal of this paper is to investigate how these techniques should be adapted to singing voice analysis, and to provide a comparative evaluation…

声音 · 计算机科学 2020-01-01 Onur Babacan , Thomas Drugman , Nicolas d'Alessandro , Nathalie Henrich , Thierry Dutoit