中文
相关论文

相关论文: A Unified Multimodal De- and Re-coupling Framework…

200 篇论文

Single modality action recognition on RGB or depth sequences has been extensively explored recently. It is generally accepted that each of these two modalities has different strengths and limitations for the task of action recognition.…

计算机视觉与模式识别 · 计算机科学 2016-12-28 Amir Shahroudy , Tian-Tsong Ng , Yihong Gong , Gang Wang

We present a novel unsupervised deep learning framework for anomalous event detection in complex video scenes. While most existing works merely use hand-crafted appearance and motion features, we propose Appearance and Motion DeepNet (AMDN)…

计算机视觉与模式识别 · 计算机科学 2015-10-07 Dan Xu , Elisa Ricci , Yan Yan , Jingkuan Song , Nicu Sebe

Current multispectral object detection methods often retain extraneous background or noise during feature fusion, limiting perceptual performance. To address this, we propose an innovative feature fusion framework based on cross-modal…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Jifeng Shen , Haibo Zhan , Xin Zuo , Heng Fan , Xiaohui Yuan , Jun Li , Wankou Yang

Robust visual tracking is a challenging computer vision problem, with many real-world applications. Most existing approaches employ hand-crafted appearance features, such as HOG or Color Names. Recently, deep RGB features extracted from…

计算机视觉与模式识别 · 计算机科学 2016-12-21 Susanna Gladh , Martin Danelljan , Fahad Shahbaz Khan , Michael Felsberg

The burgeoning volume of multi-modal data necessitates advanced retrieval paradigms beyond unimodal and cross-modal approaches. Composed Multi-modal Retrieval (CMR) emerges as a pivotal next-generation technology, enabling users to query…

In this paper, we are interested in self-supervised learning the motion cues in videos using dynamic motion filters for a better motion representation to finally boost human action recognition in particular. Thus far, the vision community…

计算机视觉与模式识别 · 计算机科学 2019-04-26 Ali Diba , Vivek Sharma , Luc Van Gool , Rainer Stiefelhagen

We propose FusionBERT, a novel multi-view visual fusion framework for image-3D multimodal retrieval. Existing image-3D representation learning methods predominantly focus on feature alignment of a single object image and its 3D model,…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Wei Li , Yufan Ren , Hanqing Jiang , Jianhui Ding , Zhen Peng , Leman Feng , Yichun Shentu , Guoqiang Xu , Baigui Sun

RGB-D scene parsing methods effectively capture both semantic and geometric features of the environment, demonstrating great potential under challenging conditions such as extreme weather and low lighting. However, existing RGB-D scene…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Jianxin Huang , Jiahang Li , Sergey Vityazev , Alexander Dvorkovich , Rui Fan

We present a new way to detect 3D objects from multimodal inputs, leveraging both LiDAR and RGB cameras in a hybrid late-cascade scheme, that combines an RGB detection network and a 3D LiDAR detector. We exploit late fusion principles to…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Carlo Sgaravatti , Roberto Basla , Riccardo Pieroni , Matteo Corno , Sergio M. Savaresi , Luca Magri , Giacomo Boracchi

Multi-modal sensor data fusion takes advantage of complementary or reinforcing information from each sensor and can boost overall performance in applications such as scene classification and target detection. This paper presents a new…

计算机视觉与模式识别 · 计算机科学 2024-02-08 Hersh Vakharia , Xiaoxiao Du

A novel deep neural network training paradigm that exploits the conjoint information in multiple heterogeneous sources is proposed. Specifically, in a RGB-D based action recognition task, it cooperatively trains a single convolutional…

计算机视觉与模式识别 · 计算机科学 2018-01-04 Pichao Wang , Wanqing Li , Jun Wan , Philip Ogunbona , Xinwang Liu

In this paper, we propose a data augmentation method for action recognition using instance segmentation. Although many data augmentation methods have been proposed for image recognition, few of them are tailored for action recognition. Our…

计算机视觉与模式识别 · 计算机科学 2022-11-15 Jun Kimata , Tomoya Nitta , Toru Tamaki

Most two-stream action recognition networks apply the same convolutional backbone to both RGB and optical flow streams, ignoring the fact that the two modalities have fundamentally different structural properties. Optical flow captures…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Md. Afzalur Rahaman , Tahmid Rahman

Action recognition has long been a fundamental and intriguing problem in artificial intelligence. The task is challenging due to the high dimensionality nature of an action, as well as the subtle motion details to be considered. Current…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Yuheng Yang , Haipeng Chen , Zhenguang Liu , Yingda Lyu , Beibei Zhang , Shuang Wu , Zhibo Wang , Kui Ren

Scene understanding based on image segmentation is a crucial component of autonomous vehicles. Pixel-wise semantic segmentation of RGB images can be advanced by exploiting complementary features from the supplementary modality (X-modality).…

计算机视觉与模式识别 · 计算机科学 2023-11-27 Jiaming Zhang , Huayao Liu , Kailun Yang , Xinxin Hu , Ruiping Liu , Rainer Stiefelhagen

Pixel space augmentation has grown in popularity in many Deep Learning areas, due to its effectiveness, simplicity, and low computational cost. Data augmentation for videos, however, still remains an under-explored research topic, as most…

计算机视觉与模式识别 · 计算机科学 2022-11-10 Artjoms Gorpincenko , Michal Mackiewicz

Indoor semantic segmentation is fundamental to computer vision and robotics, supporting applications such as autonomous navigation, augmented reality, and smart environments. Although RGB-D fusion leverages complementary appearance and…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Yan Gong , Jianli Lu , Yongsheng Gao , Jie Zhao , Xiaojuan Zhang , Susanto Rahardja

Deep learning-based image fusion approaches have obtained wide attention in recent years, achieving promising performance in terms of visual perception. However, the fusion module in the current deep learning-based methods suffers from two…

计算机视觉与模式识别 · 计算机科学 2022-02-01 Dongyu Rao , Xiao-Jun Wu , Tianyang Xu , Guoyang Chen

Multi-sensor fusion is essential for accurate 3D object detection in self-driving systems. Camera and LiDAR are the most commonly used sensors, and usually, their fusion happens at the early or late stages of 3D detectors with the help of…

计算机视觉与模式识别 · 计算机科学 2023-11-08 Javed Ahmad , Alessio Del Bue

Multi-modal object tracking integrates auxiliary modalities such as depth, thermal infrared, event flow, and language to provide additional information beyond RGB images, showing great potential in improving tracking stabilization in…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Shiyu Xuan , Zechao Li , Jinhui Tang