English
Related papers

Related papers: SIREM: Speech-Informed MRI Reconstruction with Lea…

200 papers

The training of modern speech processing systems often requires a large amount of simulated room impulse response (RIR) data in order to allow the systems to generalize well in real-world, reverberant environments. However, simulating…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-09 Yi Luo , Jianwei Yu

Based on a 3D pre-treatment magnetic resonance (MR) scan, we developed DREME-MR to jointly reconstruct the reference patient anatomy and a data-driven, patient-specific cardiorespiratory motion model. Via a motion encoder simultaneously…

Medical Physics · Physics 2025-07-04 Hua-Chieh Shao , Xiaoxue Qian , Guoping Xu , Can Wu , Ricardo Otazo , Jie Deng , You Zhang

This work proposes a self-navigated variable density spiral(VDS) based manifold regularization scheme to prospectively improve dynamic speech MRI at 3T. Short readout 1.3ms spirals were used to minimize off-resonance. A custom 16-channel…

Image and Video Processing · Electrical Eng. & Systems 2023-05-03 Rushdi Zahid Rusho , Abdul Haseeb Ahmed , Stanley Kruger , Wahidul Alam , David Meyer , David Howard , Brad Story , Mathews Jacob , Sajan Goud Lingala

We present ReCoM, an efficient framework for generating high-fidelity and generalizable human body motions synchronized with speech. The core innovation lies in the Recurrent Embedded Transformer (RET), which integrates Dynamic Embedding…

Graphics · Computer Science 2025-03-31 Yong Xie , Yunlian Sun , Hongwen Zhang , Yebin Liu , Jinhui Tang

Speech production is a complex process spanning neural planning, motor control, muscle activation, and articulatory kinematics. While the acoustic speech signal is the most accessible product of the speech production act, it does not…

We propose ARTI-6, a compact six-dimensional articulatory speech encoding framework derived from real-time MRI data that captures crucial vocal tract regions including the velum, tongue root, and larynx. ARTI-6 consists of three components:…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-27 Jihwan Lee , Sean Foley , Thanathai Lertpetchpun , Kevin Huang , Yoonjeong Lee , Tiantian Feng , Louis Goldstein , Dani Byrd , Shrikanth Narayanan

Generating radiology reports automatically reduces the workload of radiologists and helps the diagnoses of specific diseases. Many existing methods take this task as modality transfer process. However, since the key information related to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Yitian Tao , Liyan Ma , Jing Yu , Han Zhang

Purpose: To investigate whether a vision-language foundation model can enhance undersampled MRI reconstruction by providing high-level contextual information beyond conventional priors. Methods: We proposed a semantic distribution-guided…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Ruimin Feng , Xingxin He , Ronald Mercer , Zachary Stewart , Fang Liu

Magnetic Resonance Imaging (MRI) is a cornerstone in medicine and healthcare but suffers from long acquisition times. Traditional accelerated MRI methods optimize for generic image quality, lacking adaptability for specific clinical tasks.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Fangmao Ju , Yuzhu He , Zhiwen Xue , Chunfeng Lian , Jianhua Ma

We introduce SIRI, Scaling Iterative Reinforcement Learning with Interleaved Compression, a simple yet effective RL approach for Large Reasoning Models (LRMs) that enables more efficient and accurate reasoning. Existing studies have…

Machine Learning · Computer Science 2025-09-30 Haoming Wen , Yushi Bai , Juanzi Li , Jie Tang

Accurate modeling of the vocal tract is necessary to construct articulatory representations for interpretable speech processing and linguistics. However, vocal tract modeling is challenging because many internal articulators are occluded…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Rishi Jain , Bohan Yu , Peter Wu , Tejas Prabhune , Gopala Anumanchipalli

We propose SLARM, a feed-forward model that unifies dynamic scene reconstruction, semantic understanding, and real-time streaming inference. SLARM captures complex, non-uniform motion through higher-order motion modeling, trained solely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Zhicheng Qiu , Jiarui Meng , Tong-an Luo , Yican Huang , Xuan Feng , Xuanfu Li , ZHan Xu

Investigating the relationship between internal tissue point motion of the tongue and oropharyngeal muscle deformation measured from tagged MRI and intelligible speech can aid in advancing speech motor control theories and developing novel…

Image and Video Processing · Electrical Eng. & Systems 2023-02-15 Xiaofeng Liu , Fangxu Xing , Jerry L. Prince , Maureen Stone , Georges El Fakhri , Jonghye Woo

In speech enhancement (SE), phase estimation is important for perceptual quality, so many methods take clean speech's complex short-time Fourier transform (STFT) spectrum or the complex ideal ratio mask (cIRM) as the learning target. To…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-12 Yuewei Zhang , Huanbin Zou , Jie Zhu

Implicit Neural Representations (INRs) have emerged as a promising method for representing diverse data modalities, including 3D shapes, images, and audio. While recent research has demonstrated successful applications of INRs in image and…

Sound · Computer Science 2023-06-23 Luca A. Lanzendörfer , Roger Wattenhofer

Magnetic Resonance Imaging (MRI) acquisitions require extensive scan times, limiting patient throughput and increasing susceptibility to motion artifacts. Accelerated parallel MRI techniques reduce acquisition time by undersampling k-space…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Mingi Kang

Brain magnetic resonance imaging (MRI) has been extensively employed across clinical and research fields, but often exhibits sensitivity to site effects arising from non-biological variations such as differences in field strength and…

Image and Video Processing · Electrical Eng. & Systems 2024-05-31 Mengqi Wu , Lintao Zhang , Pew-Thian Yap , Hongtu Zhu , Mingxia Liu

Spatial information is a critical clue for multi-channel multi-speaker target speech recognition. Most state-of-the-art multi-channel Automatic Speech Recognition (ASR) systems extract spatial features only during the speech separation…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-27 Yiwen Shao , Yong Xu , Sanjeev Khudanpur , Dong Yu

The process of human speech production involves coordinated respiratory action to elicit acoustic speech signals. Typically, speech is produced when air is forced from the lungs and is modulated by the vocal tract, where such actions are…

Tagged magnetic resonance imaging~(MRI) has been used for decades to observe and quantify the detailed motion of deforming tissue. However, this technique faces several challenges such as tag fading, large motion, long computation times,…

Image and Video Processing · Electrical Eng. & Systems 2023-05-02 Zhangxing Bian , Fangxu Xing , Jinglun Yu , Muhan Shao , Yihao Liu , Aaron Carass , Jiachen Zhuo , Jonghye Woo , Jerry L. Prince