English
Related papers

Related papers: Prospectively accelerated dynamic speech MRI at 3 …

200 papers

Self-supervised learning via masked prediction pre-training (MPPT) has shown impressive performance on a range of speech-processing tasks. This paper proposes a method to bias self-supervised learning towards a specific task. The core idea…

Computation and Language · Computer Science 2022-11-07 Florian L. Kreyssig , Yangyang Shi , Jinxi Guo , Leda Sari , Abdelrahman Mohamed , Philip C. Woodland

Recent breakthroughs in Visual Language Models (VLMs) and Multimodal Large Language Models (MLLMs) have significantly advanced 3D scene perception towards language-driven cognition. However, existing 3D language models struggle with sparse,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Shiyu Liu , Lianlei Shan

Purpose: To develop an efficient navigator-based motion and temporal B0 shift correction technique for 3D multi-echo gradient-echo (ME-GRE) MRI for quantitative susceptibility mapping (QSM) and R2* mapping. Theory and Methods: A dual-echo…

Medical Physics · Physics 2024-03-20 Yuguang Meng , Jason W. Allen , Vahid Khalilzad Sharghi , Deqiang Qiu

Magnetic resonance imaging (MRI) is the gold standard imaging modality for numerous diagnostic tasks, yet its usefulness is tempered due to its high cost and infrastructural requirements. Low-cost very-low-field portable scanners offer new…

Pretrained vision-language models (VLMs), such as CLIP, have shown remarkable potential in few-shot image classification and led to numerous effective transfer learning strategies. These methods leverage the pretrained knowledge of VLMs to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Dexia Chen , Qianjie Zhu , Weibing Li , Yue Yu , Tong Zhang , Ruixuan Wang

Blood vessels of the brain provide the human brain with the required nutrients and oxygen. As a vulnerable part of the cerebral blood supply, pathology of small vessels can cause serious problems such as Cerebral Small Vessel Diseases…

When dealing with overlapped speech, the performance of automatic speech recognition (ASR) systems substantially degrades as they are designed for single-talker speech. To enhance ASR performance in conversational or meeting environments,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-16 Hassan Taherian , DeLiang Wang

Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recognition (ASR) for smart…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-22 Ju Lin , Niko Moritz , Yiteng Huang , Ruiming Xie , Ming Sun , Christian Fuegen , Frank Seide

Speed-of-sound is a biomechanical property for quantitative tissue differentiation, with great potential as a new ultrasound-based image modality. A conventional ultrasound array transducer can be used together with an acoustic mirror, or…

Computer Vision and Pattern Recognition · Computer Science 2018-07-20 Valery Vishnevskiy , Sergio J Sanabria , Orcun Goksel

Recent Vision Mamba (Vim) models exhibit nearly linear complexity in sequence length, making them highly attractive for processing visual data. However, the training methodologies and their potential are still not sufficiently explored. In…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Zizheng Huang , Haoxing Chen , Jiaqi Li , Jun Lan , Huijia Zhu , Weiqiang Wang , Limin Wang

Purpose: 3D Time-of-flight (TOF) MR Angiography (MRA) can accurately visualize the intracranial vasculature, but is limited by long acquisition times. Compressed sensing (CS) reconstruction can be used to substantially accelerate…

Medical Physics · Physics 2021-04-12 Matthijs H. S. de Buck , Peter Jezzard , Aaron T. Hess

We present ChiReSSD, a speech reconstruction framework that preserves children speaker's identity while suppressing mispronunciations. Unlike prior approaches trained on healthy adult speech, ChiReSSD adapts to the voices of children with…

In recent research, slight performance improvement is observed from automatic speech recognition systems to audio-visual speech recognition systems in the end-to-end framework with low-quality videos. Unmatching convergence rates and…

Computation and Language · Computer Science 2024-03-12 Yusheng Dai , Hang Chen , Jun Du , Xiaofei Ding , Ning Ding , Feijun Jiang , Chin-Hui Lee

A recent line of research on automated speaking assessment (ASA) has benefited from self-supervised learning (SSL) representations, which capture rich acoustic and linguistic patterns in non-native speech without underlying assumptions of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-23 Tien-Hong Lo , Szu-Yu Chen , Yao-Ting Sung , Berlin Chen

This paper presents novel single and multi-shell sampling schemes for diffusion MRI. In diffusion MRI, it is paramount that the number of samples is as small as possible in order that scan times are practical in a clinical setting. The…

Signal Processing · Electrical Eng. & Systems 2019-02-01 Alice P. Bates , Zubair Khalid , Jason D. McEwen , Rodney A. Kennedy , Alessandro Daducci , Erick J. Canales-Rodríguez

In this paper, we propose to extend the deep, complex U-Network architecture for speech enhancement by incorporating a probabilistic (i.e., variational) latent space model. The proposed model is evaluated against several ablated versions of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-06 Eike J. Nustede , Jörn Anemüller

Purpose: To investigate whether a vision-language foundation model can enhance undersampled MRI reconstruction by providing high-level contextual information beyond conventional priors. Methods: We proposed a semantic distribution-guided…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Ruimin Feng , Xingxin He , Ronald Mercer , Zachary Stewart , Fang Liu

Recent diffusion-based text-to-speech (TTS) models achieve high naturalness and expressiveness, yet often suffer from speaker drift, a subtle, gradual shift in perceived speaker identity within a single utterance. This underexplored…

Self-supervised speech models (S3Ms) achieve strong downstream performance, yet their learned representations remain poorly understood under natural and adversarial perturbations. Prior studies rely on representation similarity or global…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-05 Sandra Arcos-Holzinger , Sarah M. Erfani , James Bailey , Sanjeev Khudanpur

Many speech enhancement (SE) methods rely on continuous representations. Recently, discrete audio tokens have been explored to enable autoregressive generation for SE. However, it remains unclear whether discretization itself consistently…

Sound · Computer Science 2026-03-24 Jingyi Li , Luca Della Libera , Mirco Ravanelli , Cem Subakan