English
Related papers

Related papers: A real-time framework for visual feedback of artic…

200 papers

Acoustic to articulatory inversion has often been limited to a small part of the vocal tract because the data are generally EMA (ElectroMagnetic Articulography) data requiring sensors to be glued to easily accessible articulators. The…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-16 Sofiane Azzouz , Pierre-André Vuissoz , Yves Laprie

Model visualization (ModelVis) has emerged as a major research direction, yet existing taxonomies are largely organized by data or tasks, making it difficult to treat models as first-class analysis objects. We present a model-centric…

Machine Learning · Computer Science 2026-03-31 Siyu Wu , Lei Shi , Lei Xia , Cenyang Wu , Zipeng Liu , Yingchaojie Feng , Liang Zhou , Wei Chen

Integrating physiological signals such as electroencephalogram (EEG), with other data such as interview audio, may offer valuable multimodal insights into psychological states or neurological disorders. Recent advancements with Large…

Human-Computer Interaction · Computer Science 2024-08-15 Yongquan Hu , Shuning Zhang , Ting Dang , Hong Jia , Flora D. Salim , Wen Hu , Aaron J. Quigley

Remarkable effectiveness of the channel or spatial attention mechanisms for producing more discernible feature representation are illustrated in various computer vision tasks. However, modeling the cross-channel relationships with channel…

Computer Vision and Pattern Recognition · Computer Science 2023-06-07 Daliang Ouyang , Su He , Guozhong Zhang , Mingzhu Luo , Huaiyong Guo , Jian Zhan , Zhijie Huang

Multi-modal Large Language Models (MLLMs) have recently exhibited impressive general-purpose capabilities by leveraging vision foundation models to encode the core concepts of images into representations. These are then combined with…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Sara Ghazanfari , Alexandre Araujo , Prashanth Krishnamurthy , Siddharth Garg , Farshad Khorrami

Despite the recent success of Multimodal Foundation Models (FMs), their reliance on massive paired datasets limits their applicability in low-data and rare-scenario settings where aligned data is scarce and expensive. A key bottleneck is…

Machine Learning · Computer Science 2026-05-14 Truong Pham , Anay Majee , Rishabh Iyer

Multimodal learning is a rapidly growing research field that has revolutionized multitasking and generative modeling in AI. While much of the research has focused on dealing with unstructured data (e.g., language, images, audio, or video),…

Artificial Intelligence · Computer Science 2024-03-11 Marco D Alessandro , Enrique Calabrés , Mikel Elkano

Objective. In this article, we present data and methods for decoding speech articulations using surface electromyogram (EMG) signals. EMG-based speech neuroprostheses offer a promising approach for restoring audible speech in individuals…

Computation and Language · Computer Science 2025-10-07 Harshavardhana T. Gowda , Zachary D. McNaughton , Lee M. Miller

As two important textual modalities in electronic health records (EHR), both structured data (clinical codes) and unstructured data (clinical narratives) have recently been increasingly applied to the healthcare domain. Most existing…

Computation and Language · Computer Science 2022-11-01 Sicen Liu , Xiaolong Wang , Yongshuai Hou , Ge Li , Hui Wang , Hui Xu , Yang Xiang , Buzhou Tang

This article addresses the issue of representing electroencephalographic (EEG) signals in an efficient way. While classical approaches use a fixed Gabor dictionary to analyze EEG signals, this article proposes a data-driven method to obtain…

Machine Learning · Computer Science 2013-03-22 Quentin Barthélemy , Cédric Gouy-Pailler , Yoann Isaac , Antoine Souloumiac , Anthony Larue , Jérôme I. Mars

We present a multilinear statistical model of the human tongue that captures anatomical and tongue pose related shape variations separately. The model is derived from 3D magnetic resonance imaging data of 11 speakers sustaining speech…

Computer Vision and Pattern Recognition · Computer Science 2018-04-18 Alexander Hewer , Stefanie Wuhrer , Ingmar Steiner , Korin Richmond

We propose an online spatiotemporal articulation model estimation framework that estimates both articulated structure as well as a temporal prediction model solely using passive observations. The resulting model can predict future mo- tions…

Robotics · Computer Science 2016-04-13 Suren Kumar , Vikas Dhiman , Madan Ravi Ganesh , Jason J. Corso

This paper presents a novel approach towards creating a foundational model for aligning neural data and visual stimuli across multimodal representationsof brain activity by leveraging contrastive learning. We used electroencephalography…

Computer Vision and Pattern Recognition · Computer Science 2024-11-18 Matteo Ferrante , Tommaso Boccato , Grigorii Rashkov , Nicola Toschi

Multimodal learning has been proven to be an effective method to improve speech enhancement (SE) performance, especially in challenging situations such as low signal-to-noise ratios, speech noise, or unseen noise types. In previous studies,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-15 Kuan-Chen Wang , Kai-Chun Liu , Hsin-Min Wang , Yu Tsao

Streaming speech-to-avatar synthesis creates real-time animations for a virtual character from audio data. Accurate avatar representations of speech are important for the visualization of sound in linguistics, phonetics, and phonology,…

Sound · Computer Science 2023-10-26 Tejas S. Prabhune , Peter Wu , Bohan Yu , Gopala K. Anumanchipalli

The growing convergence between Large Language Models (LLMs) and electroencephalography (EEG) research is enabling new directions in neural decoding, brain-computer interfaces (BCIs), and affective computing. This survey offers a systematic…

Signal Processing · Electrical Eng. & Systems 2025-06-11 Naseem Babu , Jimson Mathew , A. P. Vinod

Acoustic articulatory inversion is a major processing challenge, with a wide range of applications from speech synthesis to feedback systems for language learning and rehabilitation. In recent years, deep learning methods have been applied…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-13 Sofiane Azzouz , Pierre-André Vuissoz , Yves Laprie

Machine learning models built on training data with multiple modalities can reveal new insights that are not accessible through unimodal datasets. For example, cardiac magnetic resonance images (MRIs) and electrocardiograms (ECGs) are both…

Machine Learning · Computer Science 2023-12-05 Bum Chul Kwon , Samuel Friedman , Kai Xu , Steven A Lubitz , Anthony Philippakis , Puneet Batra , Patrick T Ellinor , Kenney Ng

Machine learning for tabular data remains constrained by poor schema generalization, a challenge rooted in the lack of semantic understanding of structured variables. This challenge is particularly acute in domains like clinical medicine,…

Machine Learning · Computer Science 2026-05-05 Hongxi Mao , Wei Zhou , Mengting Jia , Tao Fang , Huan Gao , Bin Zhang , Shangyang Li

To build speech processing methods that can handle speech as naturally as humans, researchers have explored multiple ways of building an invertible mapping from speech to an interpretable space. The articulatory space is a promising…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-26 Peter Wu , Li-Wei Chen , Cheol Jun Cho , Shinji Watanabe , Louis Goldstein , Alan W Black , Gopala K. Anumanchipalli