中文
相关论文

相关论文: ASGIR: Audio Spectrogram Transformer Guided Classi…

200 篇论文

Automatic recognition of insect sound could help us understand changing biodiversity trends around the world -- but insect sounds are challenging to recognize even for deep learning. We present a new dataset comprised of 26399 audio files,…

声音 · 计算机科学 2025-03-20 Marius Faiß , Burooj Ghani , Dan Stowell

Across various research domains, remotely-sensed weather products are valuable for answering many scientific questions; however, their temporal and spatial resolutions are often too coarse to answer many questions. For instance, in wildlife…

音频与语音处理 · 电气工程与系统科学 2023-10-02 Enis Berk Çoban , Megan Perra , Michael I. Mandel

Music recommendation systems have emerged as a vital component to enhance user experience and satisfaction for the music streaming services, which dominates music consumption. The key challenge in improving these recommender systems lies in…

声音 · 计算机科学 2023-07-21 Junfei Zhang

Biodiversity loss poses a significant threat to humanity, making wildlife monitoring essential for assessing ecosystem health. Avian species are ideal subjects for this due to their popularity and the ease of identifying them through their…

机器学习 · 计算机科学 2026-02-23 Nina Brolich , Simon Geis , Maximilian Kasper , Alexander Barnhill , Axel Plinge , Dominik Seuß

Musical instrument classification, a key area in Music Information Retrieval, has gained considerable interest due to its applications in education, digital music production, and consumer media. Recent advances in machine learning,…

声音 · 计算机科学 2024-11-04 Joanikij Chulev

In this paper we present a research on identification of audio recording devices from background noise, thus providing a method for forensics. The audio signal is the sum of speech signal and noise signal. Usually, people pay more attention…

声音 · 计算机科学 2016-04-28 Simeng Qi , Zheng Huang , Yan Li , Shaopei Shi

This paper is about alerting acoustic event detection and sound source localisation in an urban scenario. Specifically, we are interested in spotting the presence of horns, and sirens of emergency vehicles. In order to obtain a reliable…

声音 · 计算机科学 2022-03-29 Letizia Marchegiani , Paul Newman

Acoustic scene classification systems using deep neural networks classify given recordings into pre-defined classes. In this study, we propose a novel scheme for acoustic scene classification which adopts an audio tagging system inspired by…

音频与语音处理 · 电气工程与系统科学 2020-04-21 Jee-weon Jung , Hye-jin Shim , Ju-ho Kim , Seung-bin Kim , Ha-Jin Yu

Sense of hearing is crucial for autonomous vehicles (AVs) to better perceive its surrounding environment. Although visual sensors of an AV, such as camera, lidar, and radar, help to see its surrounding environment, an AV cannot see beyond…

声音 · 计算机科学 2022-09-12 Finley Walden , Sagar Dasgupta , Mizanur Rahman , Mhafuzul Islam

Accurate estimation of aircraft operations, such as takeoffs and landings, is critical for effective airport management, yet remains challenging, especially at non-towered facilities lacking dedicated surveillance infrastructure. This paper…

声音 · 计算机科学 2025-09-15 Abdullah All Tanvir , Chenyu Huang , Moe Alahmad , Chuyang Yang , Xin Zhong

Cardiac auscultation is an essential point-of-care method used for the early diagnosis of heart diseases. Automatic analysis of heart sounds for abnormality detection is faced with the challenges of additive noise and sensor-dependent…

声音 · 计算机科学 2021-06-04 Farhat Binte Azam , Md. Istiaq Ansari , Ian Mclane , Taufiq Hasan

Automated Audio Captioning (AAC) aims to develop systems capable of describing an audio recording using a textual sentence. In contrast, Audio-Text Retrieval (ATR) systems seek to find the best matching audio recording(s) for a given…

计算与语言 · 计算机科学 2023-08-30 Etienne Labbé , Thomas Pellegrini , Julien Pinquier

Typical ASR systems segment the input audio into utterances using purely acoustic information, which may not resemble the sentence-like units that are expected by conventional machine translation (MT) systems for Spoken Language…

Passive acoustic monitoring (PAM) studies generate thousands of hours of audio, which may be used to monitor specific animal populations, conduct broad biodiversity surveys, detect threats such as poachers, and more. Machine learning…

定量方法 · 定量生物学 2024-02-26 Amanda K. Navine , Tom Denton , Matthew J. Weldy , Patrick J. Hart

Can we determine someone's geographic location purely from the sounds they hear? Are acoustic signals enough to localize within a country, state, or even city? We tackle the challenge of global-scale audio geolocation, formalize the…

声音 · 计算机科学 2025-07-23 Mustafa Chasmai , Wuao Liu , Subhransu Maji , Grant Van Horn

Automatic transcription of guitar strumming is an underrepresented and challenging task in Music Information Retrieval (MIR), particularly for extracting both strumming directions and chord progressions from audio signals. While existing…

声音 · 计算机科学 2025-08-12 Sebastian Murgul , Johannes Schimper , Michael Heizmann

Modern speech synthesis uses neural vocoders to model raw waveform samples directly. This increased versatility has expanded the scope of vocoders from speech to other domains, such as music. We address another interesting domain of…

音频与语音处理 · 电气工程与系统科学 2022-09-22 Rhythm Bhatia , Tomi H. Kinnunen

From the existing research it has been observed that many techniques and methodologies are available for performing every step of Automatic Speech Recognition (ASR) system, but the performance (Minimization of Word Error Recognition-WER and…

计算与语言 · 计算机科学 2013-03-25 Urmila Shrawankar , Vilas Thakare

Separating vocal elements from musical tracks is a longstanding challenge in audio signal processing. This study tackles the distinct separation of vocal components from musical spectrograms. We employ the Short Time Fourier Transform…

声音 · 计算机科学 2024-05-31 Adam Sorrenti