English
Related papers

Related papers: Nosey: Open-source hardware for acoustic nasalance

200 papers

This paper introduces Synthetic Enclosed Echoes (SEE), a novel dataset designed to enhance robot perception and 3D reconstruction capabilities in underwater environments. SEE comprises high-fidelity synthetic sonar data, complemented by a…

Robotics · Computer Science 2025-05-22 Guilherme de Oliveira , Matheus M. dos Santos , Paulo L. J. Drews-Jr

Our goal is to collect a large-scale audio-visual dataset with low label noise from videos in the wild using computer vision techniques. The resulting dataset can be used for training and evaluating audio recognition models. We make three…

Computer Vision and Pattern Recognition · Computer Science 2020-09-28 Honglie Chen , Weidi Xie , Andrea Vedaldi , Andrew Zisserman

Over-the-air computation has the potential to increase the communication-efficiency of data-dependent distributed wireless systems, but is vulnerable to eavesdropping. We consider over-the-air computation over block-fading additive white…

Information Theory · Computer Science 2022-12-23 Luis Maßny , Antonia Wachter-Zeh

This technical report describes the details of our TASK1A submission of the DCASE2021 challenge. The goal of the task is to design an audio scene classification system for device-imbalanced datasets under the constraints of model…

Sound · Computer Science 2022-10-26 Byeonggeun Kim , Seunghan Yang , Jangho Kim , Simyung Chang

Learning from noisy labels remains a major challenge in medical image analysis, where annotation demands expert knowledge and substantial inter-observer variability often leads to inconsistent or erroneous labels. Despite extensive research…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Yuan Ma , Junlin Hou , Chao Zhang , Yukun Zhou , Zongyuan Ge , Haoran Xie , Lie Ju

Image denoising has recently taken a leap forward due to machine learning. However, image denoisers, both expert-based and learning-based, are mostly tested on well-behaved generated noises (usually Gaussian) rather than on real-life…

Image and Video Processing · Electrical Eng. & Systems 2020-04-29 Florian Lemarchand , Eduardo Fernandes Montesuma , Maxime Pelcat , Erwan Nogues

The recommendation to change breathing patterns from the mouth to the nose can have a significantly positive impact upon the general well being of the individual. We classify nasal and mouth breathing by using an acoustic sensor and…

Neural and Evolutionary Computing · Computer Science 2010-08-26 Kevin Curran , Peng Yuan , Damian Coyle

Robotic practices on the construction site emerge as an attention-attracting manner owing to their capability of tackle complex challenges, especially in the rebar-involved scenarios. Most of existing products and research are mainly…

Robotics · Computer Science 2025-09-03 Mingze Liu , Sai Fan , Haozhen Li , Haobo Liang , Yixing Yuan , Yanke Wang

In recent years, the introduction of neural networks (NNs) into the field of speech enhancement has brought significant improvements. However, many of the proposed methods are quite demanding in terms of computational complexity and memory…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-19 Ernst Seidel , Pejman Mowlaee , Tim Fingscheidt

This paper introduces Task 2 of the DCASE2019 Challenge, titled "Audio tagging with noisy labels and minimal supervision". This task was hosted on the Kaggle platform as "Freesound Audio Tagging 2019". The task evaluates systems for…

Sound · Computer Science 2020-01-22 Eduardo Fonseca , Manoj Plakal , Frederic Font , Daniel P. W. Ellis , Xavier Serra

Data-Free Knowledge Distillation (DFKD) has made significant recent strides by transferring knowledge from a teacher neural network to a student neural network without accessing the original data. Nonetheless, existing approaches encounter…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Minh-Tuan Tran , Trung Le , Xuan-May Le , Mehrtash Harandi , Quan Hung Tran , Dinh Phung

Automatic speech recognition (ASR) degrades severely in noisy environments. Although speech enhancement (SE) front-ends effectively suppress background noise, they often introduce artifacts that harm recognition. Observation addition (OA)…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-25 Haoyang Li , Changsong Liu , Wei Rao , Hao Shi , Sakriani Sakti , Eng Siong Chng

We demonstrate tools and applications developed based on the method of "sound safeguarding," which enables any sound to be used for acoustic measurements. We developed tools for preparation, interactive and real-time measurement, and report…

Sound · Computer Science 2025-07-29 Hideki Kawahara , Kohei Yatabe , Ken-Ichi Sakakibara

Voice Conversion (VC) is a technique that aims to transform the non-linguistic information of a source utterance to change the perceived identity of the speaker. While there is a rich literature on VC, most proposed methods are trained and…

The reliability of logical operations is indispensable for the reliable operation of computational systems. Since the down-sizing of micro-fabrication generates non-negligible noise in these systems, a new approach for designing…

Other Computer Science · Computer Science 2020-04-22 Tetsuya J. Kobayashi

Optomechanical transduction is demonstrated for nanoscale torsional resonators evanescently coupled to optical microdisk whispering gallery mode resonators. The on-chip, integrated devices are measured using a fully fiber-based system,…

Mesoscale and Nanoscale Physics · Physics 2013-05-14 P. H. Kim , C. Doolin , B. D. Hauer , A. J. R. MacDonald , M. R. Freeman , P. E. Barclay , J. P. Davis

3D Gaussian Splatting (3DGS) has become one of the most promising 3D reconstruction technologies. However, label noise in real-world scenarios-such as moving objects, non-Lambertian surfaces, and shadows-often leads to reconstruction…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Han Ling , Xian Xu , Yinghui Sun , Quansen Sun

Speech synthesis might hold the key to low-resource speech recognition. Data augmentation techniques have become an essential part of modern speech recognition training. Yet, they are simple, naive, and rarely reflect real-world conditions.…

Computation and Language · Computer Science 2020-12-25 Deblin Bagchi , Shannon Wotherspoon , Zhuolin Jiang , Prasanna Muthukumar

Prosody is essential for speech technology, shaping comprehension, naturalness, and expressiveness. However, current text-to-speech (TTS) systems still struggle to accurately capture human-like prosodic variation, in part because existing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-05 Cedric Chan , Jianjing Kuang

The denoising process of diffusion models can be interpreted as an approximate projection of noisy samples onto the data manifold. Moreover, the noise level in these samples approximates their distance to the underlying manifold. Building…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Abulikemu Abuduweili , Chenyang Yuan , Changliu Liu , Frank Permenter