English
Related papers

Related papers: Multimodal active speaker detection and virtual ci…

200 papers

Speech enhancement (SE) aims to reduce noise in speech signals. Most SE techniques focus only on addressing audio information. In this work, inspired by multimodal learning, which utilizes data from different modalities, and the recent…

Sound · Computer Science 2022-04-19 Jen-Cheng Hou , Syu-Siang Wang , Ying-Hui Lai , Yu Tsao , Hsiu-Wen Chang , Hsin-Min Wang

Speech enhancement (SE) aims to reduce noise in speech signals. Most SE techniques focus only on addressing audio information. In this work, inspired by multimodal learning, which utilizes data from different modalities, and the recent…

Sound · Computer Science 2018-01-25 Jen-Cheng Hou , Syu-Siang Wang , Ying-Hui Lai , Yu Tsao , Hsiu-Wen Chang , Hsin-Min Wang

Audiovisual automatic speech recognition (AV-ASR) aims to improve the robustness of a speech recognition system by incorporating visual information. Training fully supervised multimodal models for this task from scratch, however is limited…

Computer Vision and Pattern Recognition · Computer Science 2023-03-30 Paul Hongsuck Seo , Arsha Nagrani , Cordelia Schmid

Sociometric badges are an emerging technology for study how teams interact in physical places. Audio data recorded by sociometric badges is often downsampled to not record discussions of the sociometric badges holders. To gain more…

In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at both the 7th and 8th Dialog System Technology Challenges…

Computation and Language · Computer Science 2021-10-14 Ankit P. Shah , Shijie Geng , Peng Gao , Anoop Cherian , Takaaki Hori , Tim K. Marks , Jonathan Le Roux , Chiori Hori

The rapid development of deep learning and generative AI technologies has profoundly transformed the digital contact landscape, creating realistic Deepfake that poses substantial challenges to public trust and digital media integrity. This…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Ying Xu , Marius Pedersen , Kiran Raja

Explicit communication is often valued for its directness in presenting information but requires attention during exchange, resulting in cognitive interruptions. On the other hand, implicit communication contributes to tacit and smooth…

Robotics · Computer Science 2025-10-02 Andrew Boateng , Prakhar Bhartiya , Taha Shaheen , Yu Zhang

Audio-Visual Speech Recognition (AVSR) offers a robust solution for speech recognition in challenging environments, such as cocktail-party scenarios, where relying solely on audio proves insufficient. However, current AVSR models are often…

Sound · Computer Science 2025-06-04 Thai-Binh Nguyen , Ngoc-Quan Pham , Alexander Waibel

Multi-Modal automatic speech recognition (ASR) techniques aim to leverage additional modalities to improve the performance of speech recognition systems. While existing approaches primarily focus on video or contextual information, the…

Sound · Computer Science 2023-12-27 Haoxu Wang , Fan Yu , Xian Shi , Yuezhang Wang , Shiliang Zhang , Ming Li

Recently audio-visual speech recognition (AVSR), which better leverages video modality as additional information to extend automatic speech recognition (ASR), has shown promising results in complex acoustic environments. However, there is…

Sound · Computer Science 2023-12-15 Fan Yu , Haoxu Wang , Ziyang Ma , Shiliang Zhang

This work proposes a novel method to generate realistic talking head videos using audio and visual streams. We animate a source image by transferring head motion from a driving video using a dense motion field generated using learnable…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Madhav Agarwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C V Jawahar

Face anti-spoofing (FAS) is a vital component of remote biometric authentication systems based on facial recognition, increasingly used across web-based applications. Among emerging threats, video injection attacks -- facilitated by…

Voice controlled virtual assistants (VAs) are now available in smartphones, cars, and standalone devices in homes. In most cases, the user needs to first "wake-up" the VA by saying a particular word/phrase every time he or she wants the VA…

Human-Computer Interaction · Computer Science 2019-02-05 Atta Norouzian , Bogdan Mazoure , Dermot Connolly , Daniel Willett

Recent studies have demonstrated that incorporating auxiliary information, such as speaker voiceprint or visual cues, can substantially improve Speech Enhancement (SE) performance. However, single-channel methods often yield suboptimal…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-06 Chihyun Liu , Jiaxuan Fan , Mingtung Sun , Michael Anthony , Mingsian R. Bai , Yu Tsao

Recent research has revealed that reducing the temporal and spatial redundancy are both effective approaches towards efficient video recognition, e.g., allocating the majority of computation to a task-relevant subset of frames or the most…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Yulin Wang , Yang Yue , Xinhong Xu , Ali Hassani , Victor Kulikov , Nikita Orlov , Shiji Song , Humphrey Shi , Gao Huang

Speaker-attributed automatic speech recognition (SA-ASR) in multi-party meeting scenarios is one of the most valuable and challenging ASR task. It was shown that single-channel frame-level diarization with serialized output training…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-03 Mohan Shi , Jie Zhang , Zhihao Du , Fan Yu , Qian Chen , Shiliang Zhang , Li-Rong Dai

Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-modal attention with quadratic computational cost, limiting their suitability for resource…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Qing Zhong , Guodong Ding , Lingqiao Liu , Zaiwen Feng , Lin Yuanbo Wu , Angela Yao

An immersive acoustic experience enabled by spatial audio is just as crucial as the visual aspect in creating realistic virtual environments. However, existing methods for room impulse response estimation rely either on data-demanding…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Derong Jin , Ruohan Gao

Speaker verification (SV) has recently attracted considerable research interest due to the growing popularity of virtual assistants. At the same time, there is an increasing requirement for an SV system: it should be robust to short speech…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-07 Youngmoon Jung , Yeunju Choi , Hyungjun Lim , Hoirin Kim

With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the dissonance between the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Trevine Oorloff , Surya Koppisetti , Nicolò Bonettini , Divyaraj Solanki , Ben Colman , Yaser Yacoob , Ali Shahriyari , Gaurav Bharaj
‹ Prev 1 8 9 10 Next ›