中文
相关论文

相关论文: MIR-GAN: Refining Frame-Level Modality-Invariant R…

200 篇论文

We propose a cross-modal transformer-based neural correction models that refines the output of an automatic speech recognition (ASR) system so as to exclude ASR errors. Generally, neural correction models are composed of encoder-decoder…

End-to-end models for robust automatic speech recognition (ASR) have not been sufficiently well-explored in prior work. With end-to-end models, one could choose to preprocess the input speech using speech enhancement techniques and train…

音频与语音处理 · 电气工程与系统科学 2021-02-15 Archiki Prasad , Preethi Jyothi , Rajbabu Velmurugan

Despite the rapid advance of automatic speech recognition (ASR) technologies, accurate recognition of cocktail party speech characterised by the interference from overlapping speakers, background noise and room reverberation remains a…

声音 · 计算机科学 2022-04-11 Guinan Li , Jianwei Yu , Jiajun Deng , Xunying Liu , Helen Meng

Self-supervised sound source localization is usually challenged by the modality inconsistency. In recent studies, contrastive learning based strategies have shown promising to establish such a consistent correspondence between audio and…

计算机视觉与模式识别 · 计算机科学 2023-08-10 Tianyu Liu , Peng Zhang , Wei Huang , Yufei Zha , Tao You , Yanning Zhang

We capitalize on large amounts of readily-available, synchronous data to learn a deep discriminative representations shared across three major natural modalities: vision, sound and language. By leveraging over a year of sound from video and…

计算机视觉与模式识别 · 计算机科学 2017-06-06 Yusuf Aytar , Carl Vondrick , Antonio Torralba

Recently, there has been numerous breakthroughs in face hallucination tasks. However, the task remains rather challenging in videos in comparison to the images due to inherent consistency issues. The presence of extra temporal dimension in…

计算机视觉与模式识别 · 计算机科学 2022-11-22 Shailza Sharma , Abhinav Dhall , Vinay Kumar , Vivek Singh Bawa

With the widespread application of automatic speech recognition (ASR) systems, their vulnerability to adversarial attacks has been extensively studied. However, most existing adversarial examples are generated on specific individual models,…

声音 · 计算机科学 2025-03-26 Weifei Jin , Junjie Su , Hejia Wang , Yulin Ye , Jie Hao

Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted at Interspeech 2025, which focuses on multi-modal,…

Studies on emotion recognition (ER) show that combining lexical and acoustic information results in more robust and accurate models. The majority of the studies focus on settings where both modalities are available in training and…

计算与语言 · 计算机科学 2019-06-26 Gustavo Aguilar , Viktor Rozgić , Weiran Wang , Chao Wang

There has been a recent surge in adversarial attacks on deep learning based automatic speech recognition (ASR) systems. These attacks pose new challenges to deep learning security and have raised significant concerns in deploying ASR…

密码学与安全 · 计算机科学 2021-03-08 Shehzeen Hussain , Paarth Neekhara , Shlomo Dubnov , Julian McAuley , Farinaz Koushanfar

Traditionally, research in automated speech recognition has focused on local-first encoding of audio representations to predict the spoken phonemes in an utterance. Unfortunately, approaches relying on such hyper-local information tend to…

音频与语音处理 · 电气工程与系统科学 2022-09-19 David M. Chan , Shalini Ghosh , Debmalya Chakrabarty , Björn Hoffmeister

Visual acoustic matching (VAM) is pivotal for enhancing the immersive experience, and the task of dereverberation is effective in improving audio intelligibility. Existing methods treat each task independently, overlooking the inherent…

声音 · 计算机科学 2024-07-16 Jian Ma , Wenguan Wang , Yi Yang , Feng Zheng

In this paper, we introduce a novel audio-visual multi-modal bridging framework that can utilize both audio and visual information, even with uni-modal inputs. We exploit a memory network that stores source (i.e., visual) and target (i.e.,…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Minsu Kim , Joanna Hong , Se Jin Park , Yong Man Ro

Fine-grained image search is still a challenging problem due to the difficulty in capturing subtle differences regardless of pose variations of objects from fine-grained categories. In practice, a dynamic inventory with new fine-grained…

计算机视觉与模式识别 · 计算机科学 2018-07-09 Kevin Lin , Fan Yang , Qiaosong Wang , Robinson Piramuthu

Producing a large annotated speech corpus for training ASR systems remains difficult for more than 95% of languages all over the world which are low-resourced, but collecting a relatively big unlabeled data set for such languages is more…

计算与语言 · 计算机科学 2019-08-26 Kuan-Yu Chen , Che-Ping Tsai , Da-Rong Liu , Hung-Yi Lee , Lin-shan Lee

Objective: Recognizing retinal vessel abnormity is vital to early diagnosis of ophthalmological diseases and cardiovascular events. However, segmentation results are highly influenced by elusive vessels, especially in low-contrast…

图像与视频处理 · 电气工程与系统科学 2019-12-19 Yukun Zhou , Zailiang Chen , Hailan Shen , Xianxian Zheng , Rongchang Zhao , Xuanchu Duan

Audio-Visual Speech Recognition (AVSR) leverages both acoustic and visual cues to improve speech recognition under noisy conditions. A central question is how to design a fusion mechanism that allows the model to effectively exploit visual…

音频与语音处理 · 电气工程与系统科学 2026-02-10 Seaone Ok , Min Jun Choi , Eungbeom Kim , Seungu Han , Kyogu Lee

We propose Relativistic Adversarial Feedback (RAF), a novel training objective for GAN vocoders that improves in-domain fidelity and generalization to unseen scenarios. Although modern GAN vocoders employ advanced architectures, their…

音频与语音处理 · 电气工程与系统科学 2026-03-13 Yongjoon Lee , Jung-Woo Choi

The generative adversarial networks (GANs) have facilitated the development of speech enhancement recently. Nevertheless, the performance advantage is still limited when compared with state-of-the-art models. In this paper, we propose a…

声音 · 计算机科学 2020-06-16 Andong Li , Chengshi Zheng , Renhua Peng , Cunhang Fan , Xiaodong Li

3D-aware generative adversarial networks (GANs) synthesize high-fidelity and multi-view-consistent facial images using only collections of single-view 2D imagery. Towards fine-grained control over facial attributes, recent efforts…

计算机视觉与模式识别 · 计算机科学 2023-03-14 Jingxiang Sun , Xuan Wang , Lizhen Wang , Xiaoyu Li , Yong Zhang , Hongwen Zhang , Yebin Liu