中文
相关论文

相关论文: Enabling Multi-Species Bird Classification on Low-…

200 篇论文

Voice activity detection (VAD) is the task of detecting speech in an audio stream, which is challenging due to numerous unseen noises and low signal-to-noise ratios in real environments. Recently, neural network-based VADs have alleviated…

声音 · 计算机科学 2024-05-28 Jidong Jia , Pei Zhao , Di Wang

State-of-the-art performance for many edge applications is achieved by deep neural networks (DNNs). Often, these DNNs are location- and time-sensitive, and must be delivered over a wireless channel rapidly and efficiently. In this paper, we…

网络与互联网体系结构 · 计算机科学 2023-07-21 Mikolaj Jankowski , Deniz Gunduz , Krystian Mikolajczyk

In this paper, we propose a novel Convolutional Neural Network (CNN) architecture for learning multi-scale feature representations with good tradeoffs between speed and accuracy. This is achieved by using a multi-branch network, which has…

计算机视觉与模式识别 · 计算机科学 2019-08-01 Chun-Fu Chen , Quanfu Fan , Neil Mallinar , Tom Sercu , Rogerio Feris

We present in this paper an ultra-low power (ULP) Recurrent Neural Network (RNN) based classifier for an always-on voice Wake-Up Sensor (WUS) with performances suitable for real-world applications. The purpose of our sensor is to bring down…

音频与语音处理 · 电气工程与系统科学 2021-03-09 Emmanuel Hardy , Franck Badets

DCentNet is a novel decentralized multistage signal classification approach designed for biomedical data from IoT wearable sensors, integrating early exit points (EEP) to enhance energy efficiency and processing speed. Unlike traditional…

信号处理 · 电气工程与系统科学 2025-02-26 Xiaolin Li , Binhua Huang , Barry Cardiff , Deepu John

One persistent challenge in Speech Emotion Recognition (SER) is the ubiquitous environmental noise, which frequently results in deteriorating SER performance in practice. In this paper, we introduce a Two-level Refinement Network, dubbed…

声音 · 计算机科学 2024-09-04 Chengxin Chen , Pengyuan Zhang

As unmanned aerial vehicles (UAVs) become increasingly prevalent in both consumer and defense applications, the need for reliable, modality-specific classification systems grows in urgency. This paper addresses the challenge of data…

机器学习 · 计算机科学 2025-08-15 Andrew P. Berg , Qian Zhang , Mia Y. Wang

Efficient on-device models have become attractive for near-sensor insight generation, of particular interest to the ecological conservation community. For this reason, deep learning researchers are proposing more approaches to develop lower…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Emmanuel Azuh Mensah , Joban Mand , Yueheng Ou , Min Jang , Kurtis Heimerl

Audio DeepFakes are utterances generated with the use of deep neural networks. They are highly misleading and pose a threat due to use in fake news, impersonation, or extortion. In this work, we focus on increasing accessibility to the…

声音 · 计算机科学 2022-10-13 Piotr Kawa , Marcin Plata , Piotr Syga

Ultrasound image segmentation faces unique challenges including speckle noise, low contrast, and ambiguous boundaries, while clinical deployment demands computationally efficient models. We propose USEANet, an ultrasound-specific edge-aware…

图像与视频处理 · 电气工程与系统科学 2025-09-12 Jingyi Gao , Di Wu , Baha lhnaini

Automatic modulation classification (AMC) is a crucial stage in the spectrum management, signal monitoring, and control of wireless communication systems. The accurate classification of the modulation format plays a vital role in the…

信号处理 · 电气工程与系统科学 2023-04-04 Jiawei Zhang , Tiantian Wang , Zhixi Feng , Shuyuan Yang

Precision agriculture relies heavily on effective weed management to ensure robust crop yields. This study presents RoWeeder, an innovative framework for unsupervised weed mapping that combines crop-row detection with a noise-resilient deep…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Pasquale De Marinis , Gennaro Vessio , Giovanna Castellano

Deep learning has dramatically improved the performance of sounds recognition. However, learning acoustic models directly from the raw waveform is still challenging. Current waveform-based models generally use time-domain convolutional…

声音 · 计算机科学 2018-03-29 Boqing Zhu , Changjian Wang , Feng Liu , Jin Lei , Zengquan Lu , Yuxing Peng

This study revisits the findings of Carl et al., who evaluated the pre-trained Google Inception-ResNet-v2 model for automated detection of European wild mammal species in camera trap images. To assess the reproducibility and…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Tobias Abraham Haider

Beamforming has been extensively investigated for multi-channel audio processing tasks. Recently, learning-based beamforming methods, sometimes called \textit{neural beamformers}, have achieved significant improvements in both signal…

音频与语音处理 · 电气工程与系统科学 2019-10-02 Yi Luo , Enea Ceolini , Cong Han , Shih-Chii Liu , Nima Mesgarani

This paper presents a low cost, on premise system for autonomous backyard bird monitoring in Belgian urban gardens. A motion triggered IP camera uploads short clips via FTP to a local server, where frames are sampled and birds are localized…

计算机视觉与模式识别 · 计算机科学 2025-08-14 El Mustapha Mansouri

We present SVCnet, a system for modelling speaker variability. Encoder Neural Networks specialized for each speech sound produce low dimensionality models of acoustical variation, and these models are further combined into an overall model…

声音 · 计算机科学 2022-11-17 Michael Witbrock , Patrick Haffner

Bird strikes pose a significant threat to aviation safety, often resulting in loss of life, severe aircraft damage, and substantial financial costs. Existing bird strike prevention strategies primarily rely on avian radar systems that…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Elaheh Sabziyan Varnousfaderani , Syed A. M. Shihab , Jonathan King

The persisting threats on migratory bird populations highlight the urgent need for effective monitoring techniques that could assist in their conservation. Among these, passive acoustic monitoring is an essential tool, particularly for…

声音 · 计算机科学 2025-05-26 Louis Airale , Adrien Pajot , Juliette Linossier

In this work, we derive a generic overcomplete frame thresholding scheme based on risk minimization. Overcomplete frames being favored for analysis tasks such as classification, regression or anomaly detection, we provide a way to leverage…

音频与语音处理 · 电气工程与系统科学 2017-12-27 Romain Cosentino , Randall Balestriero , Richard Baraniuk , Ankit Patel