中文
相关论文

相关论文: ASGIR: Audio Spectrogram Transformer Guided Classi…

200 篇论文

Accurate estimation of Room Impulse Response (RIR), which captures an environment's acoustic properties, is important for speech processing and AR/VR applications. We propose AV-RIR, a novel multi-modal multi-task learning approach to…

声音 · 计算机科学 2024-04-25 Anton Ratnarajah , Sreyan Ghosh , Sonal Kumar , Purva Chiniya , Dinesh Manocha

Environmental Sound Classification (ESC) is an active research area in the audio domain and has seen a lot of progress in the past years. However, many of the existing approaches achieve high accuracy by relying on domain-specific features…

计算机视觉与模式识别 · 计算机科学 2020-04-17 Andrey Guzhov , Federico Raue , Jörn Hees , Andreas Dengel

Animals hear and vocalize across frequency ranges that differ substantially from humans, often extending into the ultrasonic domain. Yet most computational bioacoustics systems rely on audio models pre-trained at 16 kHz, restricting their…

The capability for environmental sound recognition (ESR) can determine the fitness of individuals in a way to avoid dangers or pursue opportunities when critical sound events occur. It still remains mysterious about the fundamental…

神经与进化计算 · 计算机科学 2019-02-05 Qiang Yu , Yanli Yao , Longbiao Wang , Huajin Tang , Jianwu Dang , Kay Chen Tan

Acoustic environment characterization opens doors for sound reproduction innovations, smart EQing, speech enhancement, hearing aids, and forensics. Reverberation time, clarity, and direct-to-reverberant ratio are acoustic parameters that…

声音 · 计算机科学 2020-10-22 Paul Callens , Milos Cernak

Mastering fine-grained visual recognition, essential in many expert domains, can require that specialists undergo years of dedicated training. Modeling the progression of such expertize in humans remains challenging, and accurately…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Leonie Bossemeyer , Samuel Heinrich , Grant Van Horn , Oisin Mac Aodha

Acoustic recognition has emerged as a prominent task in deep learning research, frequently utilizing spectral feature extraction techniques such as the spectrogram from the Short-Time Fourier Transform and the scalogram from the Wavelet…

音频与语音处理 · 电气工程与系统科学 2025-12-01 Dang Thoai Phan

Bird strikes present a huge risk for aircraft, especially since traditional airport bird surveillance is mainly dependent on inefficient human observation. Computer vision based technology has been proposed to automatically detect birds,…

计算机视觉与模式识别 · 计算机科学 2016-01-19 Ying Huang , Hong Zheng , Haibin Ling , Erik Blasch , Hao Yang

Automatic Speech Recognition (ASR) systems must be robust to the myriad types of noises present in real-world environments including environmental noise, room impulse response, special effects as well as attacks by malicious actors…

声音 · 计算机科学 2024-09-26 Muhammad A. Shah , Bhiksha Raj

Air traffic management and specifically air-traffic control (ATC) rely mostly on voice communications between Air Traffic Controllers (ATCos) and pilots. In most cases, these voice communications follow a well-defined grammar that could be…

Passive acoustic monitoring (PAM) is crucial for bioacoustic research, enabling non-invasive species tracking and biodiversity monitoring. Citizen science platforms provide large annotated datasets from focal recordings, where the target…

声音 · 计算机科学 2026-01-29 Ilyass Moummad , Romain Serizel , Emmanouil Benetos , Nicolas Farrugia

We present BEAMER: a new spatially exploitative approach to learning object detectors which shows excellent results when applied to the task of detecting objects in greyscale aerial imagery in the presence of ambiguous and noisy data. There…

计算机视觉与模式识别 · 计算机科学 2009-07-27 Damian Eads , Edward Rosten , David Helmbold

In the context of the Internet of Things (IoT), sound sensing applications are required to run on embedded platforms where notions of product pricing and form factor impose hard constraints on the available computing power. Whereas…

声音 · 计算机科学 2016-09-09 Siddharth Sigtia , Adam M. Stark , Sacha Krstulovic , Mark D. Plumbley

In this work, we derive a generic overcomplete frame thresholding scheme based on risk minimization. Overcomplete frames being favored for analysis tasks such as classification, regression or anomaly detection, we provide a way to leverage…

音频与语音处理 · 电气工程与系统科学 2017-12-27 Romain Cosentino , Randall Balestriero , Richard Baraniuk , Ankit Patel

Accurately detecting voiced intervals in speech signals is a critical step in pitch tracking and has numerous applications. While conventional signal processing methods and deep learning algorithms have been proposed for this task, their…

音频与语音处理 · 电气工程与系统科学 2023-12-07 Yixuan Zhang , Heming Wang , DeLiang Wang

Audio fingerprinting techniques have seen great advances in recent years, enabling accurate and fast audio retrieval even in conditions when the queried audio sample has been highly deteriorated or recorded in noisy conditions. Expectedly,…

信息检索 · 计算机科学 2025-09-26 Kemal Altwlkany , Sead Delalić , Adis Alihodžić , Elmedin Selmanović , Damir Hasić

The goal of this project is to classify four different insect sounds: cicada, beetle, termite, and cricket. One application of this project is for pest control to monitor and protect our ecosystem. Our project leverages data augmentation,…

声音 · 计算机科学 2024-12-18 Yinxuan Wang , Sudip Vhaduri

Recent years have seen immense progress in 3D computer vision and computer graphics, with emerging tools that can virtualize real-world 3D environments for numerous Mixed Reality (XR) applications. However, alongside immersive visual…

声音 · 计算机科学 2024-06-12 Mason Wang , Ryosuke Sawata , Samuel Clarke , Ruohan Gao , Shangzhe Wu , Jiajun Wu

We introduce a novel approach to studying animal behaviour and the context in which it occurs, through the use of microphone backpacks carried on the backs of individual free-flying birds. These sensors are increasingly used by animal…

声音 · 计算机科学 2016-12-19 Dan Stowell , Emmanouil Benetos , Lisa F. Gill

Anomalous sound detection (ASD) in the wild requires robustness to distribution shifts such as unseen low-SNR input mixtures of machine and noise types. State-of-the-art systems extract embeddings from an adapted audio encoder and detect…

音频与语音处理 · 电气工程与系统科学 2025-10-30 Phurich Saengthong , Tomoya Nishida , Kota Dohi , Natsuo Yamashita , Yohei Kawaguchi
‹ 上一页 1 8 9 10 下一页 ›