English
Related papers

Related papers: Audio-Visual Event Localization via Recursive Fusi…

200 papers

This paper presents a novel vehicle motion forecasting method based on multi-head attention. It produces joint forecasts for all vehicles on a road scene as sequences of multi-modal probability density functions of their positions. Its…

Machine Learning · Computer Science 2019-12-23 Jean Mercat , Thomas Gilles , Nicole El Zoghby , Guillaume Sandou , Dominique Beauvois , Guillermo Pita Gil

Phase-based features related to vocal source characteristics can be incorporated into magnitude-based speaker recognition systems to improve the system performance. However, traditional feature-level fusion methods typically ignore the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-20 Rongfeng Su , Mengjie Du , Xiaokang Liu , Lan Wang , Nan Yan

Inspired by the complementarity between conventional frame-based and bio-inspired event-based cameras, we propose a multi-modal based approach to fuse visual cues from the frame- and event-domain to enhance the single object tracking…

Computer Vision and Pattern Recognition · Computer Science 2021-09-21 Jiqing Zhang , Xin Yang , Yingkai Fu , Xiaopeng Wei , Baocai Yin , Bo Dong

Audio-visual navigation represents a significant area of research in which intelligent agents utilize egocentric visual and auditory perceptions to identify audio targets. Conventional navigation methodologies typically adopt a staged…

Artificial Intelligence · Computer Science 2025-10-01 Hailong Zhang , Yinfeng Yu , Liejun Wang , Fuchun Sun , Wendong Zheng

Various attention mechanisms are being widely applied to acoustic scene classification. However, we empirically found that the attention mechanism can excessively discard potentially valuable information, despite improving performance. We…

Machine Learning · Computer Science 2021-12-24 Hye-jin Shim , Jee-weon Jung , Ju-ho Kim , Ha-Jin Yu

In this work, we explore the impact of visual modality in addition to speech and text for improving the accuracy of the emotion detection system. The traditional approaches tackle this task by fusing the knowledge from the various…

Machine Learning · Computer Science 2020-04-24 Seunghyun Yoon , Subhadeep Dey , Hwanhee Lee , Kyomin Jung

In this work, we present a deep learning-based approach for image tampering localization fusion. This approach is designed to combine the outcomes of multiple image forensics algorithms and provides a fused tampering localization map, which…

Computer Vision and Pattern Recognition · Computer Science 2021-05-14 Polychronis Charitidis , Giorgos Kordopatis-Zilos , Symeon Papadopoulos , Ioannis Kompatsiaris

In this paper, we propose two techniques, namely joint modeling and data augmentation, to improve system performances for audio-visual scene classification (AVSC). We employ pre-trained networks trained only on image data sets to extract…

Learning social media content is the basis of many real-world applications, including information retrieval and recommendation systems, among others. In contrast with previous works that focus mainly on single modal or bi-modal learning, we…

Computation and Language · Computer Science 2021-03-24 Hongru Liang , Haozheng Wang , Jun Wang , Shaodi You , Zhe Sun , Jin-Mao Wei , Zhenglu Yang

Learning common subspace is prevalent way in cross-modal retrieval to solve the problem of data from different modalities having inconsistent distributions and representations that cannot be directly compared. Previous cross-modal retrieval…

Multimedia · Computer Science 2021-10-27 Donghuo Zeng , Jianming Wu , Gen Hattori , Yi Yu , Rong Xu

To understand and quantify the quality of mixed-presence collaboration around wall-sized displays, robust evaluation methodologies are needed, that are adapted for a room-sized experience and are not perceived as obtrusive. In this paper,…

Human-Computer Interaction · Computer Science 2025-07-22 Adrien Coppens , Valérie Maquil

In recent years, various applications in computer vision have achieved substantial progress based on deep learning, which has been widely used for image fusion and shown to achieve adequate performance. However, suffering from limited…

Computer Vision and Pattern Recognition · Computer Science 2022-08-16 Zhengwen Shen , Jun Wang , Zaiyu Pan , Yulian Li , Jiangyu Wang

In this paper, we introduce a novel audio-visual multi-modal bridging framework that can utilize both audio and visual information, even with uni-modal inputs. We exploit a memory network that stores source (i.e., visual) and target (i.e.,…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Minsu Kim , Joanna Hong , Se Jin Park , Yong Man Ro

Acoustic event detection and scene classification are major research tasks in environmental sound analysis, and many methods based on neural networks have been proposed. Conventional methods have addressed these tasks separately; however,…

Multiple object tracking (MOT) is a significant task in achieving autonomous driving. Traditional works attempt to complete this task, either based on point clouds (PC) collected by LiDAR, or based on images captured from cameras. However,…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Guangming Wang , Chensheng Peng , Jinpeng Zhang , Hesheng Wang

Effective fusion of data from multiple modalities, such as video, speech, and text, is challenging due to the heterogeneous nature of multimodal data. In this paper, we propose adaptive fusion techniques that aim to model context from…

Computation and Language · Computer Science 2021-01-27 Gaurav Sahu , Olga Vechtomova

Leveraging temporal synchronization and association within sight and sound is an essential step towards robust localization of sounding objects. To this end, we propose a space-time memory network for sounding object localization in videos.…

Computer Vision and Pattern Recognition · Computer Science 2021-11-11 Sizhe Li , Yapeng Tian , Chenliang Xu

A major challenge for video captioning is to combine audio and visual cues. Existing multi-modal fusion methods have shown encouraging results in video understanding. However, the temporal structures of multiple modalities at different…

Computation and Language · Computer Science 2018-04-17 Xin Wang , Yuan-Fang Wang , William Yang Wang

Infrared and visible image fusion (IVIF) is a fundamental task in multi-modal perception that aims to integrate complementary structural and textural cues from different spectral domains. In this paper, we propose FusionNet, a novel…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Tianyao Sun , Dawei Xiang , Tianqi Ding , Xiang Fang , Yijiashun Qi , Zunduo Zhao

Vision Transformer (ViT) self-attention mechanism is characterized by feature collapse in deeper layers, resulting in the vanishing of low-level visual features. However, such features can be helpful to accurately represent and identify…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Anxhelo Diko , Danilo Avola , Marco Cascio , Luigi Cinque
‹ Prev 1 4 5 6 7 8 10 Next ›