中文
相关论文

相关论文: STNet: Deep Audio-Visual Fusion Network for Robust…

200 篇论文

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target…

计算与语言 · 计算机科学 2025-11-17 Tuochao Chen , Bandhav Veluri , Hongyu Gong , Shyamnath Gollakota

Audio-visual Navigation refers to an agent utilizing visual and auditory information in complex 3D environments to accomplish target localization and path planning, thereby achieving autonomous navigation. The core challenge of this task…

声音 · 计算机科学 2026-04-06 Xinyu Zhou , Yinfeng Yu

In multichannel speech enhancement, both spectral and spatial information are vital for discriminating between speech and noise. How to fully exploit these two types of information and their temporal dynamics remains an interesting research…

音频与语音处理 · 电气工程与系统科学 2022-11-17 Yujie Yang , Changsheng Quan , Xiaofei Li

Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on precomputed room impulse responses (RIRs) for binaural audio…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Yichen Zeng , Hebaixu Wang , Meng Liu , Yu Zhou , Chen Gao , Kehan Chen , Gongping Huang

This paper introduces the task of Auditory Referring Multi-Object Tracking (AR-MOT), which dynamically tracks specific objects in a video sequence based on audio expressions and appears as a challenging problem in autonomous driving. Due to…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Jiacheng Lin , Jiajun Chen , Kunyu Peng , Xuan He , Zhiyong Li , Rainer Stiefelhagen , Kailun Yang

We propose an audio-visual spatial-temporal deep neural network with: (1) a visual block containing a pretrained 2D-CNN followed by a temporal convolutional network (TCN); (2) an aural block containing several parallel TCNs; and (3) a…

计算机视觉与模式识别 · 计算机科学 2021-08-18 Su Zhang , Yi Ding , Ziquan Wei , Cuntai Guan

By implicitly recognizing a user based on his/her speech input, speaker identification enables many downstream applications, such as personalized system behavior and expedited shopping checkouts. Based on whether the speech content is…

机器学习 · 计算机科学 2021-06-21 Ruirui Li , Chelsea J. -T. Ju , Zeya Chen , Hongda Mao , Oguz Elibol , Andreas Stolcke

Existing works on weakly-supervised audio-visual video parsing adopt hybrid attention network (HAN) as the multi-modal embedding to capture the cross-modal context. It embeds the audio and visual modalities with a shared network, where the…

计算机视觉与模式识别 · 计算机科学 2023-11-15 Yating Xu , Conghui Hu , Gim Hee Lee

Vehicle location prediction or vehicle tracking is a significant topic within connected vehicles. This task, however, is difficult if only a single modal data is available, probably causing bias and impeding the accuracy. With the…

计算机视觉与模式识别 · 计算机科学 2018-11-08 Yue Zhang , Bin Song , Xiaojiang Du , Mohsen Guizani

This paper investigates how to perform robust visual tracking in adverse and challenging conditions using complementary visual and thermal infrared data (RGBT tracking). We propose a novel deep network architecture called qualityaware…

计算机视觉与模式识别 · 计算机科学 2019-10-15 Yabin Zhu , Chenglong Li , Bin Luo , Jin Tang

Active speaker detection (ASD) in egocentric videos presents unique challenges due to unstable viewpoints, motion blur, and off-screen speech sources - conditions under which traditional visual-centric methods degrade significantly. We…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Yu Wang , Juhyung Ha , David J. Crandall

Deepfakes are synthetic media generated using deep generative algorithms and have posed a severe societal and political threat. Apart from facial manipulation and synthetic voice, recently, a novel kind of deepfakes has emerged with either…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Vinaya Sree Katamneni , Ajita Rattani

Scene depth information can help visual information for more accurate semantic segmentation. However, how to effectively integrate multi-modality information into representative features is still an open problem. Most of the existing work…

计算机视觉与模式识别 · 计算机科学 2021-05-11 Yuejiao Su , Yuan Yuan , Zhiyu Jiang

The ability to learn robust multi-modality representation has played a critical role in the development of RGBT tracking. However, the regular fusion paradigm and the invariable tracking template remain restrictive to the feature…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Ruichao Hou , Boyue Xu , Tongwei Ren , Gangshan Wu

Audio-visual speech enhancement system is regarded as one of promising solutions for isolating and enhancing speech of desired speaker. Typical methods focus on predicting clean speech spectrum via a naive convolution neural network based…

音频与语音处理 · 电气工程与系统科学 2022-07-01 Xinmeng Xu , Yang Wang , Jie Jia , Binbin Chen , Dejun Li

In audio-visual navigation (AVN) tasks, an embodied agent must autonomously localize a sound source in unknown and complex 3D environments based on audio-visual signals. Existing methods often rely on static modality fusion strategies and…

人工智能 · 计算机科学 2025-09-23 Jia Li , Yinfeng Yu , Liejun Wang , Fuchun Sun , Wendong Zheng

This paper presents a self-supervised method for visual detection of the active speaker in a multi-person spoken interaction scenario. Active speaker detection is a fundamental prerequisite for any artificial cognitive system attempting to…

计算机视觉与模式识别 · 计算机科学 2019-07-19 Kalin Stefanov , Jonas Beskow , Giampiero Salvi

State-of-the-art Active Speaker Detection (ASD) approaches mainly use audio and facial features as input. However, the main hypothesis in this paper is that body dynamics is also highly correlated to "speaking" (and "listening") actions and…

计算机视觉与模式识别 · 计算机科学 2024-12-12 Tiago Roxo , Joana C. Costa , Pedro Inácio , Hugo Proença

Audio-visual navigation combines sight and hearing to navigate to a sound-emitting source in an unmapped environment. While recent approaches have demonstrated the benefits of audio input to detect and find the goal, they focus on clean and…

声音 · 计算机科学 2023-01-04 Abdelrahman Younes , Daniel Honerkamp , Tim Welschehold , Abhinav Valada

In the past, Acoustic Scene Classification systems have been based on hand crafting audio features that are input to a classifier. Nowadays, the common trend is to adopt data driven techniques, e.g., deep learning, where audio…

声音 · 计算机科学 2018-06-29 Eduardo Fonseca , Rong Gong , Xavier Serra