中文
相关论文

相关论文: Enabling Multi-Species Bird Classification on Low-…

200 篇论文

Speech, Music and Noise classification/segmentation is an important preprocessing step for audio processing/indexing. To this end, we propose a novel 1D Convolutional Neural Network (CNN) - SwishNet. It is a fast and lightweight…

机器学习 · 计算机科学 2018-12-04 Md. Shamim Hussain , Mohammad Ariful Haque

Sensor nodes in a wireless sensor network (WSN) for security surveillance applications should preferably be small, energy-efficient, and inexpensive with in-sensor computational abilities. An appropriate data processing scheme in the sensor…

Most work in audio enhancement targets human speech, while bioacoustics is less studied due to noisy recordings and the distinct traits of animal sounds. To fill this gap, we adapt speech enhancement methods and build BioSEN, a model made…

声音 · 计算机科学 2026-05-15 Tianyu Song , Ton Viet Ta , Ngamta Thamwattana , Hisako Nomura , Linh Thi Hoai Nguyen

Selective weed treatment is a critical step in autonomous crop management as related to crop health and yield. However, a key challenge is reliable, and accurate weed detection to minimize damage to surrounding plants. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2017-09-12 Inkyu Sa , Zetao Chen , Marija Popovic , Raghav Khanna , Frank Liebisch , Juan Nieto , Roland Siegwart

Automatic Music Transcription (AMT) has been recognized as a key enabling technology with a wide range of applications. Given the task's complexity, best results have typically been reported for systems focusing on specific settings, e.g.…

The ongoing biodiversity crisis, driven by factors such as land-use change and global warming, emphasizes the need for effective ecological monitoring methods. Acoustic monitoring of biodiversity has emerged as an important monitoring tool.…

声音 · 计算机科学 2023-12-18 Drew Priebe , Burooj Ghani , Dan Stowell

Human voice is the source of several important information. This is in the form of features. These Features help in interpreting various features associated with the speaker and speech. The speaker dependent work researchersare targeted…

声音 · 计算机科学 2022-03-30 Shankhanil Ghosh , Chhanda Saha , Naagamani Molakathaala

Automated bioacoustic analysis aids understanding and protection of both marine and terrestrial animals and their habitats across extensive spatiotemporal scales, and typically involves analyzing vast collections of acoustic data. With the…

音频与语音处理 · 电气工程与系统科学 2023-12-22 Burooj Ghani , Tom Denton , Stefan Kahl , Holger Klinck

We present working notes on transfer learning with semi-supervised dataset annotation for the BirdCLEF 2023 competition, focused on identifying African bird species in recorded soundscapes. Our approach utilizes existing off-the-shelf…

声音 · 计算机科学 2024-07-10 Anthony Miyaguchi , Nathan Zhong , Murilo Gustineli , Chris Hayduk

The drone has been used for various purposes, including military applications, aerial photography, and pesticide spraying. However, the drone is vulnerable to external disturbances, and malfunction in propellers and motors can easily occur.…

声音 · 计算机科学 2023-04-25 Wonjun Yi , Jung-Woo Choi , Jae-Woo Lee

Recurrent neural networks (RNNs) have shown promising results in audio and speech processing applications due to their strong capabilities in modelling sequential data. In many applications, RNNs tend to outperform conventional models based…

密码学与安全 · 计算机科学 2017-09-25 Jagmohan Chauhan , Suranga Seneviratne , Yining Hu , Archan Misra , Aruna Seneviratne , Youngki Lee

The BirdCLEF+ 2025 challenge requires classifying 206 species, including birds, mammals, insects, and amphibians, from soundscape recordings under a strict 90-minute CPU-only inference deadline, making many state-of-the-art deep learning…

声音 · 计算机科学 2025-07-14 Anthony Miyaguchi , Murilo Gustineli , Adrian Cheung

Deploying deep learning models in agriculture is difficult because edge devices have limited resources, but this work presents a compressed version of EcoWeedNet using structured channel pruning, quantization-aware training (QAT), and…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Omar H. Khater , Abdul Jabbar Siddiqui , Aiman El-Maleh , M. Shamim Hossain

In many real-world datasets, like WebVision, the performance of DNN based classifier is often limited by the noisy labeled data. To tackle this problem, some image related side information, such as captions and tags, often reveal underlying…

计算机视觉与模式识别 · 计算机科学 2020-09-07 Lele Cheng , Xiangzeng Zhou , Liming Zhao , Dangwei Li , Hong Shang , Yun Zheng , Pan Pan , Yinghui Xu

Weakly Supervised Semantic Segmentation (WSSS) employs weak supervision, such as image-level labels, to train the segmentation model. Despite the impressive achievement in recent WSSS methods, we identify that introducing weak labels with…

计算机视觉与模式识别 · 计算机科学 2024-07-01 Junsung Park , Hyunjung Shim

For the task of subdecimeter aerial imagery segmentation, fine-grained semantic segmentation results are usually difficult to obtain because of complex remote sensing content and optical conditions. Recently, convolutional neural networks…

计算机视觉与模式识别 · 计算机科学 2018-08-28 Kai Yue , Lei Yang , Ruirui Li , Wei Hu , Fan Zhang , Wei Li

This paper proposes a data-efficient, semi-supervised, two-pass framework for segmenting bird vocalizations. The framework utilizes a binary classification model to categorize frames of an input audio recording into the background or bird…

音频与语音处理 · 电气工程与系统科学 2019-02-27 Anshul Thakur , Padmanabhan Rajan

This research addresses the pressing challenge of enhancing processing times and detection capabilities in Unmanned Aerial Vehicle (UAV)/drone imagery for global wildfire detection, despite limited datasets. Proposing a Segmented Neural…

计算机视觉与模式识别 · 计算机科学 2024-05-02 Aditya V. Jonnalagadda , Hashim A. Hashim

Passive acoustic monitoring enables large-scale biodiversity assessment, but reliable classification of bioacoustic sounds requires not only high accuracy but also well-calibrated uncertainty estimates to ground decision-making. In…

In speaker verification systems, the utilization of short utterances presents a persistent challenge, leading to performance degradation primarily due to insufficient phonetic information to characterize the speakers. To overcome this…

音频与语音处理 · 电气工程与系统科学 2024-06-12 Seung-bin Kim , Chan-yeong Lim , Jungwoo Heo , Ju-ho Kim , Hyun-seo Shin , Kyo-Won Koo , Ha-Jin Yu