English
Related papers

Related papers: You Only Hear Once: A YOLO-like Algorithm for Audi…

200 papers

Most sound event detection (SED) systems perform well on clean datasets but degrade significantly in noisy environments. Language-queried audio source separation (LASS) models show promise for robust SED by separating target events;…

Sound · Computer Science 2025-08-12 Yuanjian Chen , Yang Xiao , Han Yin , Yadong Guan , Xubo Liu

We present AERO, a audio super-resolution model that processes speech and music signals in the spectral domain. AERO is based on an encoder-decoder architecture with U-Net like skip connections. We optimize the model using both time and…

Sound · Computer Science 2023-02-28 Moshe Mandel , Or Tal , Yossi Adi

For visually impaired people, it is highly difficult to make independent movement and safely move in both indoors and outdoors environment. Furthermore, these physically and visually challenges prevent them from in day-today live…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Heba Najm , Khirallah Elferjani , Alhaam Alariyibi

Sound Event Detection (SED) aims to predict the temporal boundaries of all the events of interest and their class labels, given an unconstrained audio sample. Taking either the splitand-classify (i.e., frame-level) strategy or the more…

Sound · Computer Science 2023-08-21 Swapnil Bhosale , Sauradip Nag , Diptesh Kanojia , Jiankang Deng , Xiatian Zhu

Recent Large Audio-Language Models (LALMs) exhibit impressive capabilities in understanding audio content for conversational QA tasks. However, these models struggle to accurately understand timestamps for temporal localization (e.g.,…

Sound · Computer Science 2025-12-15 Hualei Wang , Yiming Li , Shuo Ma , Hong Liu , Xiangdong Wang

Detecting small to tiny targets in infrared images is a challenging task in computer vision, especially when it comes to differentiating these targets from noisy or textured backgrounds. Traditional object detection methods such as YOLO…

Computer Vision and Pattern Recognition · Computer Science 2024-02-06 Alina Ciocarlan , Sylvie Le Hégarat-Mascle , Sidonie Lefebvre , Arnaud Woiselle , Clara Barbanson

A key function of auditory cognition is the association of characteristic sounds with their corresponding semantics over time. Humans attempting to discriminate between fine-grained audio categories, often replay the same discriminative…

Sound · Computer Science 2023-03-14 Alexandros Stergiou , Dima Damen

Sound event detection is the task of recognizing sounds and determining their extent (onset/offset times) within an audio clip. Existing systems commonly predict sound presence confidence in short time frames. Then, thresholding produces…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-07 Janek Ebbers , Francois G. Germain , Gordon Wichern , Jonathan Le Roux

Audio fingerprinting, also named as audio hashing, has been well-known as a powerful technique to perform audio identification and synchronization. It basically involves two major steps: fingerprint (voice pattern) design and matching…

Sound · Computer Science 2015-02-25 Ngoc Q. K. Duong , Hien-Thanh Duong

Underwater object detection (UOD) remains a critical challenge in computer vision due to underwater distortions which degrade low-level features and compromise the reliability of even state-of-the-art detectors. While YOLO models have…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Edwine Nabahirwa , Wei Song , Minghua Zhang , Shufan Chen

DNN/Accelerator co-design has shown great potential in improving QoR and performance. Typical approaches separate the design flow into two-stage: (1) designing an application-specific DNN model with high accuracy; (2) building an…

Signal Processing · Electrical Eng. & Systems 2020-05-15 Weiwei Chen , Ying Wang , Shuang Yang , Chen Liu , Lei Zhang

We introduce You Only Train Once (YOTO), a dynamic human generation framework, which performs free-viewpoint rendering of different human identities with distinct motions, via only one-time training from monocular videos. Most prior works…

Computer Vision and Pattern Recognition · Computer Science 2023-03-13 Jaehyeok Kim , Dongyoon Wee , Dan Xu

The audio-visual segmentation (AVS) task aims to segment sounding objects from a given video. Existing works mainly focus on fusing audio and visual features of a given video to achieve sounding object masks. However, we observed that prior…

Sound · Computer Science 2023-08-02 Chen Liu , Peike Li , Xingqun Qi , Hu Zhang , Lincheng Li , Dadong Wang , Xin Yu

It is generally assumed that number of classes is fixed in current audio classification methods, and the model can recognize pregiven classes only. When new classes emerge, the model needs to be retrained with adequate samples of all…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-06 Yanxiong Li , Wenchang Cao , Jialong Li , Wei Xie , Qianhua He

Environment Sound Classification has been a well-studied research problem in the field of signal processing and up till now more focus has been laid on fully supervised approaches. Over the last few years, focus has moved towards…

Object detection and segmentation represents the basis for many tasks in computer and machine vision. In biometric recognition systems the detection of the region-of-interest (ROI) is one of the most crucial steps in the overall processing…

Computer Vision and Pattern Recognition · Computer Science 2019-02-04 Žiga Emeršič , Luka Lan Gabriel , Vitomir Štruc , Peter Peer

Existing multimodal audio generation models often lack precise user control, which limits their applicability in professional Foley workflows. In particular, these models focus on the entire video and do not provide precise methods for…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Ilpo Viertola , Vladimir Iashin , Esa Rahtu

You Only Look Once (YOLO) has been the prominent model for computer vision in deep learning for a decade. This study explores the novel aspects of YOLO26, the most recent version in the YOLO series. The elimination of Distribution Focal…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Priyanto Hidayatullah , Refdinal Tubagus

This study conducted a comprehensive performance evaluation on YOLO11 (or YOLOv11) and YOLOv8, the latest in the "You Only Look Once" (YOLO) series, focusing on their instance segmentation capabilities for immature green apples in orchard…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Ranjan Sapkota , Manoj Karkee

Machine learning is a promising technique for angle-of-arrival (AOA) estimation of waves impinging a sensor array. However, the majority of the methods proposed so far only consider a known, fixed number of impinging waves, i.e., a fixed…

Signal Processing · Electrical Eng. & Systems 2021-11-19 Noud Kanters , Andrés Alayón Glazunov