English
Related papers

Related papers: PEAF: Learnable Power Efficient Analog Acoustic Fe…

200 papers

Compared with automatic speech recognition (ASR), the human auditory system is more adept at handling noise-adverse situations, including environmental noise and channel distortion. To mimic this adeptness, auditory models have been widely…

Computation and Language · Computer Science 2016-09-16 Peng Dai , Xue Teng , Frank Rudzicz , Ing Yann Soon

Audio Foundation Models (AFMs), a specialized category of Generative AI (GenAI), have the potential to transform signal processing (SP) education by integrating core applications such as speech and audio enhancement, denoising, source…

Signal Processing · Electrical Eng. & Systems 2026-04-16 Muhammad Salman Khan , Ahmad Ullah , Siddique Latif , Junaid Qadir

Nowadays, speech is becoming a more common, if not standard, interface to technology. This can be seen in the trend of technology changes over the years. Increasingly, voice is used to control programs, appliances and personal devices…

Human-Computer Interaction · Computer Science 2019-09-10 Abraham Glasser

In many speech-enabled human-machine interaction scenarios, user speech can overlap with the device playback audio. In these instances, the performance of tasks such as keyword-spotting (KWS) and device-directed speech detection (DDD) can…

Sound · Computer Science 2022-10-05 Samuele Cornell , Thomas Balestri , Thibaud Sénéchal

Mel-filterbanks are fixed, engineered audio features which emulate human perception and have been used through the history of audio understanding up to today. However, their undeniable qualities are counterbalanced by the fundamental…

Sound · Computer Science 2021-01-22 Neil Zeghidour , Olivier Teboul , Félix de Chaumont Quitry , Marco Tagliasacchi

The Keyword Spotting (KWS) task involves continuous audio stream monitoring to detect predefined words, requiring low energy devices for continuous processing. Neuromorphic devices effectively address this energy challenge. However, the…

Neural and Evolutionary Computing · Computer Science 2024-08-12 Sidi Yaya Arnaud Yarga , Sean U. N. Wood

Spoofed audio, i.e. audio that is manipulated or AI-generated deepfake audio, is difficult to detect when only using acoustic features. Some recent innovative work involving AI-spoofed audio detection models augmented with phonetic and…

Sound · Computer Science 2024-10-22 Zahra Khanjani , Christine Mallinson , James Foulds , Vandana P Janeja

We propose a novel approach for spoofed speech characterization through explainable probabilistic attribute embeddings. In contrast to high-dimensional raw embeddings extracted from a spoofing countermeasure (CM) whose dimensions are not…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-18 Manasi Chhibber , Jagabandhu Mishra , Hyejin Shim , Tomi H. Kinnunen

Active speaker detection is a challenging task in audio-visual scenario understanding, which aims to detect who is speaking in one or more speakers scenarios. This task has received extensive attention as it is crucial in applications such…

Computer Vision and Pattern Recognition · Computer Science 2023-03-09 Junhua Liao , Haihan Duan , Kanghui Feng , Wanbing Zhao , Yanbing Yang , Liangyin Chen

Neural network models for audio tasks, such as automatic speech recognition (ASR) and acoustic scene classification (ASC), are susceptible to noise contamination for real-life applications. To improve audio quality, an enhancement module,…

Using audio and text embeddings jointly for Keyword Spotting (KWS) has shown high-quality results, but the key challenge of how to semantically align two embeddings for multi-word keywords of different sequence lengths remains largely…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-09 Kumari Nishu , Minsik Cho , Devang Naik

Passive Acoustic Monitoring (PAM) is an efficient and non-invasive method for surveying ecosystems at a reduced cost. Typically, autonomous recorders allow the acquisition of vast bioacoustic datasets which are then analyzed. However, power…

Sound · Computer Science 2026-05-06 Louis Lerbourg , Paul Peyret , Juliette Linossier , Marielle Malfante

Wake word (WW) spotting is challenging in far-field due to the complexities and variations in acoustic conditions and the environmental interference in signal transmission. A suite of carefully designed and optimized audio front-end (AFE)…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-15 Yixin Gao , Noah D. Stein , Chieh-Chi Kao , Yunliang Cai , Ming Sun , Tao Zhang , Shiv Vitaladevuni

Automatic speaker verification, like every other biometric system, is vulnerable to spoofing attacks. Using only a few minutes of recorded voice of a genuine client of a speaker verification system, attackers can develop a variety of…

Sound · Computer Science 2019-06-20 Balamurali BT , Kin Wah Edward Lin , Simon Lui , Jer-Ming Chen , Dorien Herremans

In this paper, for real time enhancement of noisy speech, a method of threshold determination based on modeling of Teager energy (TE) operated perceptual wavelet packet (PWP) coefficients of the noisy speech and noise by an Erlang-2 PDF is…

Audio and Speech Processing · Electrical Eng. & Systems 2018-02-13 Md Tauhidul Islam , Celia Shahnaz , Wei-Ping Zhu , M. Omair Ahmad

Edge-based automatic speech recognition (ASR) technologies are increasingly prevalent in the development of intelligent and personalized assistants. However, resource-constrained ASR models face significant challenges in adaptivity,…

Computation and Language · Computer Science 2024-12-24 Amir Nassereldine , Dancheng Liu , Chenhui Xu , Ruiyang Qin , Yiyu Shi , Jinjun Xiong

Recent studies have introduced methods for learning acoustic word embeddings (AWEs)---fixed-size vector representations of words which encode their acoustic features. Despite the widespread use of AWEs in speech processing research, they…

Computation and Language · Computer Science 2020-04-06 Yevgen Matusevych , Herman Kamper , Sharon Goldwater

Analog signal processing (ASP) is presented as a systematic approach to address future challenges in high speed and high frequency microwave applications. The general concept of ASP is explained with the help of examples emphasizing basic…

Optics · Physics 2013-08-06 Christophe Caloz , Shulabh Gupta , Qingfeng Zhang , Babak Nikfal

Anti-spoofing is the task of speech authentication. That is, identifying genuine human speech compared to spoofed speech. The main focus of this paper is to suggest new representations for genuine and spoofed speech, based on the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-28 Matan Karo , Arie Yeredor , Itshak Lapidot

Despite rapid advancement in recent years, current speech enhancement models often produce speech that differs in perceptual quality from real clean speech. We propose a learning objective that formalizes differences in perceptual quality,…