中文
相关论文

相关论文: Prospectively accelerated dynamic speech MRI at 3 …

200 篇论文

This paper investigates the use of target-speaker automatic speech recognition (TS-ASR) for simultaneous speech recognition and speaker diarization of single-channel dialogue recordings. TS-ASR is a technique to automatically extract and…

计算与语言 · 计算机科学 2019-09-19 Naoyuki Kanda , Shota Horiguchi , Yusuke Fujita , Yawen Xue , Kenji Nagamatsu , Shinji Watanabe

In this work, we propose a deep learning approach for parallel magnetic resonance imaging (MRI) reconstruction, termed a variable splitting network (VS-Net), for an efficient, high-quality reconstruction of undersampled multi-coil MR data.…

图像与视频处理 · 电气工程与系统科学 2019-07-24 Jinming Duan , Jo Schlemper , Chen Qin , Cheng Ouyang , Wenjia Bai , Carlo Biffi , Ghalib Bello , Ben Statton , Declan P O'Regan , Daniel Rueckert

The purpose of this study is to enable in-vivo three-dimensional (3-D) ultrasound localization microscopy (ULM) of posterior ocular microvasculature using a 256-channel system and a 1024-element matrix array, and to overcome limitations of…

This paper introduces a discrete diffusion model (DDM) framework for text-aligned speech tokenization and reconstruction. By replacing the auto-regressive speech decoder with a discrete diffusion counterpart, our model achieves…

音频与语音处理 · 电气工程与系统科学 2025-09-25 Pin-Jui Ku , He Huang , Jean-Marie Lemercier , Subham Sekhar Sahoo , Zhehuai Chen , Ante Jukić

Open-vocabulary semantic segmentation (OVSS) underpins many vision and robotics tasks that require generalizable semantic understanding. Existing approaches either rely on limited segmentation training data, which hinders generalization, or…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Omar Alama , Darshil Jariwala , Avigyan Bhattacharya , Seungchan Kim , Wenshan Wang , Sebastian Scherer

In this study, we propose to investigate triplet loss for the purpose of an alternative feature representation for ASR. We consider a general non-semantic speech representation, which is trained with a self-supervised criteria based on…

声音 · 计算机科学 2021-09-24 Szu-Jui Chen , Wei Xia , John H. L. Hansen

Visual Speech Recognition (VSR) aims to infer speech into text depending on lip movements alone. As it focuses on visual information to model the speech, its performance is inherently sensitive to personal lip appearances and movements, and…

计算与语言 · 计算机科学 2024-10-21 Minsu Kim , Hyung-Il Kim , Yong Man Ro

Functional MRI (fMRI) is an important tool for non-invasive studies of brain function. Over the past decade, multi-echo fMRI methods that sample multiple echo times has become popular with potential to improve quantification. While these…

图像与视频处理 · 电气工程与系统科学 2023-12-12 Hongyi Gu , Chi Zhang , Zidan Yu , Christoph Rettenmeier , V. Andrew Stenger , Mehmet Akçakaya

Purpose: To introduce a novel deep learning based approach for fast and high-quality dynamic multi-coil MR reconstruction by learning a complementary time-frequency domain network that exploits spatio-temporal correlations simultaneously…

图像与视频处理 · 电气工程与系统科学 2021-06-21 Chen Qin , Jinming Duan , Kerstin Hammernik , Jo Schlemper , Thomas Küstner , René Botnar , Claudia Prieto , Anthony N. Price , Joseph V. Hajnal , Daniel Rueckert

Audio-visual speech recognition (AVSR) combines audio-visual modalities to improve speech recognition, especially in noisy environments. However, most existing methods deploy the unidirectional enhancement or symmetric fusion manner, which…

多媒体 · 计算机科学 2025-08-12 Junxiao Xue , Xiaozhen Liu , Xuecheng Wu , Xinyi Yin , Danlei Huang , Fei Yu

Several successful iterative approaches have recently been proposed for parallel-imaging reconstructions of variable-density (VD) acquisitions, but they often induce substantial computational burden for non-Cartesian data. Here we propose a…

医学物理 · 物理学 2017-02-07 Yiğit Baran Can , Efe Ilıcak , Tolga Çukur

Traditional clinical approaches for assessing nasality, such as nasopharyngoscopy and nasometry, involve unpleasant experiences and are problematic for children. Speech Inversion (SI), a noninvasive technique, offers a promising alternative…

音频与语音处理 · 电气工程与系统科学 2025-09-12 Saba Tabatabaee , Suzanne Boyce , Liran Oren , Mark Tiede , Carol Espy-Wilson

Audio-visual speech recognition (AVSR) can effectively and significantly improve the recognition rates of small-vocabulary systems, compared to their audio-only counterparts. For large-vocabulary systems, however, there are still many…

音频与语音处理 · 电气工程与系统科学 2021-09-13 Wentao Yu , Steffen Zeiler , Dorothea Kolossa

Supervised Deep-Learning (DL)-based reconstruction algorithms have shown state-of-the-art results for highly-undersampled dynamic Magnetic Resonance Imaging (MRI) reconstruction. However, the requirement of excessive high-quality…

图像与视频处理 · 电气工程与系统科学 2025-11-11 Jie Feng , Ruimin Feng , Qing Wu , Zhiyong Zhang , Yuyao Zhang , Hongjiang Wei

Accurately reconstructing dense and semantically annotated 3D meshes from monocular images remains a challenging task due to the lack of geometry guidance and imperfect view-dependent 2D priors. Though we have witnessed recent advancements…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Zhenhua Du , Binbin Xu , Haoyu Zhang , Kai Huo , Shuaifeng Zhi

As the popularity of deep learning (DL) in the field of magnetic resonance imaging (MRI) continues to rise, recent research has indicated that DL-based MRI reconstruction models might be excessively sensitive to minor input disturbances,…

图像与视频处理 · 电气工程与系统科学 2025-10-07 Shijun Liang , Van Hoang Minh Nguyen , Jinghan Jia , Ismail Alkhouri , Sijia Liu , Saiprasad Ravishankar

Multimodal speech recognition aims to improve the performance of automatic speech recognition (ASR) systems by leveraging additional visual information that is usually associated to the audio input. While previous approaches make crucial…

声音 · 计算机科学 2022-04-29 Dan Oneata , Horia Cucu

Fake speech detection systems have become a necessity to combat against speech deepfakes. Current systems exhibit poor generalizability on out-of-domain speech samples due to lack to diverse training data. In this paper, we attempt to…

音频与语音处理 · 电气工程与系统科学 2025-08-27 Rishith Sadashiv T N , Abhishek Bedge , Saisha Suresh Bore , Jagabandhu Mishra , Mrinmoy Bhattacharjee , S R Mahadeva Prasanna

We propose an approach for dense semantic 3D reconstruction which uses a data term that is defined as potentials over viewing rays, combined with continuous surface area penalization. Our formulation is a convex relaxation which we augment…

计算机视觉与模式识别 · 计算机科学 2019-08-27 Nikolay Savinov , Christian Haene , Lubor Ladicky , Marc Pollefeys

This paper proposes VARA-TTS, a non-autoregressive (non-AR) text-to-speech (TTS) model using a very deep Variational Autoencoder (VDVAE) with Residual Attention mechanism, which refines the textual-to-acoustic alignment layer-wisely.…

声音 · 计算机科学 2021-02-15 Peng Liu , Yuewen Cao , Songxiang Liu , Na Hu , Guangzhi Li , Chao Weng , Dan Su