English
Related papers

Related papers: End-to-end audio-visual learning for cochlear impl…

200 papers

This paper presents, a first of its kind, audio-visual (AV) speech enhacement challenge in real-noisy settings. A detailed description of the AV challenge, a novel real noisy AV corpus (ASPIRE), benchmark speech enhancement task, and…

Sound · Computer Science 2019-10-02 Mandar Gogate , Ahsan Adeel , Kia Dashtipour , Peter Derleth , Amir Hussain

Tens of millions of people live blind, and their number is ever increasing. Visual-to-auditory sensory substitution (SS) encompasses a family of cheap, generic solutions to assist the visually impaired by conveying visual information…

Neurons and Cognition · Quantitative Biology 2019-07-16 Viktor Tóth , Lauri Parkkonen

Acoustic word embeddings (AWEs) aims to map a variable-length speech segment into a fixed-dimensional representation. High-quality AWEs should be invariant to variations, such as duration, pitch and speaker. In this paper, we introduce a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-20 Jingru Lin , Xianghu Yue , Junyi Ao , Haizhou Li

The field of speech recognition is in the midst of a paradigm shift: end-to-end neural networks are challenging the dominance of hidden Markov models as a core technology. Using an attention mechanism in a recurrent encoder-decoder…

Sound · Computer Science 2017-03-16 Tsubasa Ochiai , Shinji Watanabe , Takaaki Hori , John R. Hershey

The cardiac dipole has been shown to propagate to the ears, now a common site for consumer wearable electronics, enabling the recording of electrocardiogram (ECG) signals. However, in-ear ECG recordings often suffer from significant noise…

Objective: In cochlear implant users with residual acoustic hearing, compound action potentials (CAPs) can be evoked by acoustic (aCAP) or electric (eCAP) stimulation and recorded through the electrodes of the implant. We propose a novel…

Medical Physics · Physics 2024-10-28 Daniel Kipping , Yixuan Zhang , Waldo Nogueira

Recently, a generative variational autoencoder (VAE) has been proposed for speech enhancement to model speech statistics. However, this approach only uses clean speech in the training phase, making the estimation particularly sensitive to…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-18 Huajian Fang , Guillaume Carbajal , Stefan Wermter , Timo Gerkmann

Over 1.5 billion people worldwide live with hearing impairment. Despite various technologies that have been created for individuals with such disabilities, most of these technologies are either extremely expensive or inaccessible for…

Sound · Computer Science 2023-07-11 Jesse Choe , Siddhant Sood , Ryan Park

Speech enhancement (SE) methods mainly focus on recovering clean speech from noisy input. In real-world speech communication, however, noises often exist in not only speaker but also listener environments. Although SE methods can suppress…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-23 Haoyu Li , Yun Liu , Junichi Yamagishi

Echo and noise suppression is an integral part of a full-duplex communication system. Many recent acoustic echo cancellation (AEC) systems rely on a separate adaptive filtering module for linear echo suppression and a neural module for…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-07 Karn N. Watcharasupat , Thi Ngoc Tho Nguyen , Woon-Seng Gan , Shengkui Zhao , Bin Ma

Deep learning has brought impressive progress in the study of both automatic speaker verification (ASV) and spoofing countermeasures (CM). Although solutions are mutually dependent, they have typically evolved as standalone sub-systems…

Speech Emotion Recognition (SER) often operates on speech segments detected by a Voice Activity Detection (VAD) model. However, VAD models may output flawed speech segments, especially in noisy environments, resulting in degraded…

Sound · Computer Science 2024-10-18 Natsuo Yamashita , Masaaki Yamamoto , Yohei Kawaguchi

Speech decoding from EEG signals is a challenging task, where brain activity is modeled to estimate salient characteristics of acoustic stimuli. We propose FESDE, a novel framework for Fully-End-to-end Speech Decoding from EEG signals. Our…

Signal Processing · Electrical Eng. & Systems 2024-06-14 Jihwan Lee , Aditya Kommineni , Tiantian Feng , Kleanthis Avramidis , Xuan Shi , Sudarsana Kadiri , Shrikanth Narayanan

Despite recent advances in voice separation methods, many challenges remain in realistic scenarios such as noisy recording and the limits of available data. In this work, we propose to explicitly incorporate the phonetic and linguistic…

Existing deep learning (DL) based speech enhancement approaches are generally optimised to minimise the distance between clean and enhanced speech features. These often result in improved speech quality however they suffer from a lack of…

Sound · Computer Science 2021-11-19 Tassadaq Hussain , Mandar Gogate , Kia Dashtipour , Amir Hussain

Audio-visual speech enhancement (AV-SE) aims to enhance degraded speech along with extra visual information such as lip videos, and has been shown to be more effective than audio-only speech enhancement. This paper proposes further…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-21 Rui-Chen Zheng , Yang Ai , Zhen-Hua Ling

End-to-end models for robust automatic speech recognition (ASR) have not been sufficiently well-explored in prior work. With end-to-end models, one could choose to preprocess the input speech using speech enhancement techniques and train…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-15 Archiki Prasad , Preethi Jyothi , Rajbabu Velmurugan

The integration of artificial intelligence into hearing assistance marks a paradigm shift from traditional amplification-based systems to intelligent, context-aware audio processing. This systematic literature review evaluates advances in…

Sound · Computer Science 2025-08-05 Haris Khan , Shumaila Asif , Hassan Nasir , Kamran Aziz Bhatti , Shahzad Amin Sheikh

Brain-computer interface (BCI) is the technology that enables the communication between humans and devices by reflecting status and intentions of humans. When conducting imagined speech, the users imagine the pronunciation as if actually…

Human-Computer Interaction · Computer Science 2021-12-15 Dae-Hyeok Lee , Sung-Jin Kim , Keon-Woo Lee

Acoustic echo cancellation (AEC) aims to remove interference signals while leaving near-end speech least distorted. As the indistinguishable patterns between near-end speech and interference signals, near-end speech can't be separated…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-27 Chang Han , Xinmeng Xu , Weiping Tu , Yuhong Yang , Yajie Liu