English
Related papers

Related papers: Blind Normalization of Speech From Different Chann…

200 papers

A novel method to solve inverse problems for the wave equation is introduced. The method is a combination of the boundary control method and an iterative time reversal scheme, leading to adaptive imaging of coefficient functions of the wave…

Analysis of PDEs · Mathematics 2007-05-23 Kenrick Bingham , Yaroslav Kurylev , Matti Lassas , Samuli Siltanen

This paper shows how a machine, which observes stimuli through an uncharacterized, uncalibrated channel and sensor, can glean machine-independent information (i.e., channel- and sensor-independent information) about the stimuli. First, we…

Computer Vision and Pattern Recognition · Computer Science 2009-11-10 David N. Levin

The intelligibility of speech relies on the ability of interlocutors to dynamically align their expectations about the rates at which informative changes in signals occur. Exactly how this is achieved remains an open question. We propose…

Neurons and Cognition · Quantitative Biology 2022-11-03 Maja Linke , Michael Ramscar

Machine learning techniques have garnered great interest in designing communication systems owing to their capacity in tackling with channel uncertainty. To provide theoretical guarantees for learning-based communication systems, some…

Machine Learning · Computer Science 2025-06-17 Zheshun Wu , Junfan Li , Zenglin Xu , Sumei Sun , Jie Liu

Speaker embeddings extracted with deep 2D convolutional neural networks are typically modeled as projections of first and second order statistics of channel-frequency pairs onto a linear layer, using either average or attentive pooling…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-08 Themos Stafylakis , Johan Rohdin , Lukas Burget

Shannon's theory of zero-error communication is re-examined in the broader setting of using one classical channel to simulate another exactly, and in the presence of various resources that are all classes of non-signalling correlations:…

Quantum Physics · Physics 2016-11-17 Toby S. Cubitt , Debbie Leung , William Matthews , Andreas Winter

Silent speech interfaces have been recently proposed as a way to enable communication when the acoustic signal is not available. This introduces the need to build visual speech recognition systems for silent and whispered speech. However,…

Computer Vision and Pattern Recognition · Computer Science 2018-02-20 Stavros Petridis , Jie Shen , Doruk Cetin , Maja Pantic

When recorded in an enclosed room, a sound signal will most certainly get affected by reverberation. This not only undermines audio quality, but also poses a problem for many human-machine interaction technologies that use speech as their…

Sound · Computer Science 2018-09-21 Francisco Ibarrola , Leandro Di Persia , Ruben Spies

The interdependence and high dimensionality of multivariate signals present significant challenges for denoising, as conventional univariate methods often struggle to capture the complex interactions between variables. A successful approach…

Machine Learning · Computer Science 2024-07-29 Jaesung Choi , Pilwon Kim

Wavefield focusing is often achieved by Time-Reversal Mirrors, where wavefields emitted by a source located at the focal point are evaluated at a closed boundary and sent back, after Time-Reversal, into the medium from that boundary.…

Classical Physics · Physics 2025-09-12 Giovanni Angelo Meles , Joost van der Neut , Koen W. A. van Dongen , Kees Wapenaar

Speech emotion conversion aims to convert the expressed emotion of a spoken utterance to a target emotion while preserving the lexical information and the speaker's identity. In this work, we specifically focus on in-the-wild emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-06 Navin Raj Prabhu , Nale Lehmann-Willenbrock , Timo Gerkmann

Informed speaker extraction aims to extract a target speech signal from a mixture of sources given prior knowledge about the desired speaker. Recent deep learning-based methods leverage a speaker discriminative model that maps a reference…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-17 Mohamed Elminshawi , Wolfgang Mack , Emanuël A. P. Habets

We present a visually grounded model of speech perception which projects spoken utterances and images to a joint semantic space. We use a multi-layer recurrent highway network to model the temporal nature of spoken speech, and show that it…

Computation and Language · Computer Science 2018-10-30 Grzegorz Chrupała , Lieke Gelderloos , Afra Alishahi

Time series forecasting traditionally relies on unimodal numerical inputs, which often struggle to capture high-level semantic patterns due to their dense and unstructured nature. While recent approaches have explored representing time…

Machine Learning · Computer Science 2025-07-02 Sixun Dong , Wei Fan , Teresa Wu , Yanjie Fu

Speech restoration in real-world conditions is challenging due to compounded distortions and mismatches between input and desired output rates. Most existing systems assume a fixed and shared input-output rate, relying on external…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-29 Ui-Hyeop Shin , Jaehyun Ko , Woocheol Jeong , Hyung-Min Park

This paper introduces a novel technique for reconstructing the phase of modified spectrograms of audio signals. From the analysis of mixtures of sinusoids we obtain relationships between phases of successive time frames in the…

Sound · Computer Science 2016-05-25 Paul Magron , Roland Badeau , Bertrand David

We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representations for speech content, prosodic information, and speaker…

In this paper, we propose an effective training strategy to ex-tract robust speaker representations from a speech signal. Oneof the key challenges in speaker recognition tasks is to learnlatent representations or embeddings containing…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-05 Yoohwan Kwon , Soo-Whan Chung , Hong-Goo Kang

We propose a mathematical theory for the refocusing properties observed in time-reversal experiments, where classical waves propagate through a medium, are recorded in time, then time-reversed and sent back into the medium. The salient…

Mesoscale and Nanoscale Physics · Physics 2007-05-23 Guillaume Bal , Leonid Ryzhik

In the domain of unsupervised learning most work on speech has focused on discovering low-level constructs such as phoneme inventories or word-like units. In contrast, for written language, where there is a large body of work on…

Computation and Language · Computer Science 2018-10-29 Grzegorz Chrupała , Lieke Gelderloos , Ákos Kádár , Afra Alishahi