English
Related papers

Related papers: Coding Speech through Vocal Tract Kinematics

200 papers

Understanding cognitive processes in the brain demands sophisticated models capable of replicating neural dynamics at large scales. We present a physiologically inspired speech recognition architecture, compatible and scalable with deep…

Computation and Language · Computer Science 2024-09-26 Alexandre Bittar , Philip N. Garner

Decoding speech directly from neural activity is a central goal in brain-computer interface (BCI) research. In recent years, exciting advances have been made through the growing use of intracranial field potential recordings, such as…

Signal Processing · Electrical Eng. & Systems 2025-05-28 Hui Zheng , Hai-Teng Wang , Yi-Tao Jing , Pei-Yang Lin , Han-Qing Zhao , Wei Chen , Peng-Hu Wei , Yong-Zhi Shan , Guo-Guang Zhao , Yun-Zhe Liu

Previous speech pre-training methods, such as wav2vec2.0 and HuBERT, pre-train a Transformer encoder to learn deep representations from audio data, with objectives predicting either elements from latent vector quantized space or…

Sound · Computer Science 2022-04-08 Shuo Ren , Shujie Liu , Yu Wu , Long Zhou , Furu Wei

In recent years, Speech Emotion Recognition (SER) has been investigated mainly transforming the speech signal into spectrograms that are then classified using Convolutional Neural Networks pretrained on generic images and fine tuned with…

Sound · Computer Science 2022-11-07 A. Arezzo , S. Berretti

We introduce BANC, a neural binaural audio codec designed for efficient speech compression in single and two-speaker scenarios while preserving the spatial location information of each speaker. Our key contributions are as follows: 1) The…

Sound · Computer Science 2024-11-26 Anton Ratnarajah , Shi-Xiong Zhang , Dong Yu

Acoustic vowel dynamics have some speaker-identifying characteristics, which have been ascribed to individual properties of articulatory strategies: formant transitions have a particular shape because speakers move their articulators, using…

Computation and Language · Computer Science 2026-05-25 Patrycja Strycharczuk , Justin J. H. Lo , Sam Kirkham

Voice Conversion (VC) modifies speech to match a target speaker while preserving linguistic content. Traditional methods usually extract speaker information directly from speech while neglecting the explicit utilization of linguistic…

Multimedia · Computer Science 2025-06-04 Fengjin Li , Jie Wang , Yadong Niu , Yongqing Wang , Meng Meng , Jian Luan , Zhiyong Wu

Audio-visual automatic speech recognition (AV-ASR) is an extension of ASR that incorporates visual cues, often from the movements of a speaker's mouth. Unlike works that simply focus on the lip motion, we investigate the contribution of…

Computer Vision and Pattern Recognition · Computer Science 2022-06-16 Valentin Gabeur , Paul Hongsuck Seo , Arsha Nagrani , Chen Sun , Karteek Alahari , Cordelia Schmid

The rapid population aging has stimulated the development of assistive devices that provide personalized medical support to the needies suffering from various etiologies. One prominent clinical application is a computer-assisted speech…

Computation and Language · Computer Science 2019-05-22 Emre Yılmaz , Vikramjit Mitra , Ganesh Sivaraman , Horacio Franco

The subcortical sensory pathways are the fundamental channels for mapping the outside world to our minds. Sensory pathways efficiently transmit information by adapting neural responses to the local statistics of the sensory input. The…

Neurons and Cognition · Quantitative Biology 2020-03-26 Alejandro Tabas , Glad Mihai , Stefan Kiebel , Robert Trampel , Katharina von Kriegstein

In this work, we investigate the joint use of articulatory and acoustic features for automatic speech recognition (ASR) of pathological speech. Despite long-lasting efforts to build speaker- and text-independent ASR systems for people with…

Computation and Language · Computer Science 2018-07-31 Emre Yılmaz , Vikramjit Mitra , Chris Bartels , Horacio Franco

Voice cloning is the task of learning to synthesize the voice of an unseen speaker from a few samples. While current voice cloning methods achieve promising results in Text-to-Speech (TTS) synthesis for a new voice, these approaches lack…

Sound · Computer Science 2021-02-02 Paarth Neekhara , Shehzeen Hussain , Shlomo Dubnov , Farinaz Koushanfar , Julian McAuley

Deep learning dominates speech processing but relies on massive datasets, global backpropagation-guided weight updates, and produces entangled representations. Assembly Calculus (AC), which models sparse neuronal assemblies via Hebbian…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-19 Trevor Adelson , Vidhyasaharan Sethu , Ting Dang

The way speakers articulate is well known to be variable across individuals while at the same time subject to anatomical and biomechanical constraints. In this study, we ask whether articulatory strategy in vowel production can be…

Computation and Language · Computer Science 2025-05-28 Justin J. H. Lo , Patrycja Strycharczuk , Sam Kirkham

Speech emotion recognition (SER), the task of identifying the expression of emotion from spoken content, is challenging due to the difficulty in extracting representations that capture emotional attributes from speech. The scarcity of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-27 Soumya Dutta , Sriram Ganapathy

Complex-valued sparse coding is a data representation which employs a dictionary of two-dimensional subspaces, while imposing a sparse, factorial prior on complex amplitudes. When trained on a dataset of natural image patches, it learns…

Machine Learning · Computer Science 2014-02-19 Wiktor Mlynarski

Representing speech as discrete tokens provides a framework for transforming speech into a format that closely resembles text, thus enabling the use of speech as an input to the widely successful large language models (LLMs). Currently,…

Active speaker detection (ASD) in videos with multiple speakers is a challenging task as it requires learning effective audiovisual features and spatial-temporal correlations over long temporal windows. In this paper, we present SPELL, a…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Kyle Min , Sourya Roy , Subarna Tripathi , Tanaya Guha , Somdeb Majumdar

Recent self-supervised learning (SSL) models have proven to learn rich representations of speech, which can readily be utilized by diverse downstream tasks. To understand such utilities, various analyses have been done for speech SSL models…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-24 Cheol Jun Cho , Peter Wu , Abdelrahman Mohamed , Gopala K. Anumanchipalli

Speech codecs serve as bridges between continuous speech signals and large language models, yet face an inherent conflict between acoustic fidelity and semantic preservation. To mitigate this conflict, prevailing methods augment acoustic…

Sound · Computer Science 2026-01-28 Xin Zhang , Lin Li , Xiangni Lu , Jianquan Liu , Kong Aik Lee