中文
相关论文

相关论文: Multi-Modality Fusion based on Consensus-Voting an…

200 篇论文

Existing Multi-view Clustering (MVC) methods based on subspace learning focus on consensus representation learning while neglecting the inherent topological structure of data. Despite the integration of Graph Neural Networks (GNNs) into…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Chenping Pei , Fadi Dornaika , Jingjun Bi

3D spatial information is known to be beneficial to the semantic segmentation task. Most existing methods take 3D spatial data as an additional input, leading to a two-stream segmentation network that processes RGB and 3D spatial…

计算机视觉与模式识别 · 计算机科学 2021-02-24 Lin-Zhuo Chen , Zheng Lin , Ziqin Wang , Yong-Liang Yang , Ming-Ming Cheng

Nowadays, hand gesture recognition has become an alternative for human-machine interaction. It has covered a large area of applications like 3D game technology, sign language interpreting, VR (virtual reality) environment, and robotics. But…

计算机视觉与模式识别 · 计算机科学 2022-02-28 Abir Sen , Tapas Kumar Mishra , Ratnakar Dash

Deep convolutional neural networks (ConvNets) have been recently shown to attain state-of-the-art performance for action recognition on standard-resolution videos. However, less attention has been paid to recognition performance at…

计算机视觉与模式识别 · 计算机科学 2018-10-09 Jiawei Chen , Jonathan Wu , Janusz Konrad , Prakash Ishwar

Efficiently exploiting multi-modal inputs for accurate RGB-D saliency detection is a topic of high interest. Most existing works leverage cross-modal interactions to fuse the two streams of RGB-D for intermediate features' enhancement. In…

计算机视觉与模式识别 · 计算机科学 2022-08-31 Zongwei Wu , Shriarulmozhivarman Gobichettipalayam , Brahim Tamadazte , Guillaume Allibert , Danda Pani Paudel , Cédric Demonceaux

We tackle the problem of using 3D information in convolutional neural networks for down-stream recognition tasks. Using depth as an additional channel alongside the RGB input has the scale variance problem present in image convolution based…

计算机视觉与模式识别 · 计算机科学 2018-12-05 Hang Chu , Wei-Chiu Ma , Kaustav Kundu , Raquel Urtasun , Sanja Fidler

In this paper, we address the 3D object detection task by capturing multi-level contextual information with the self-attention mechanism and multi-scale feature fusion. Most existing 3D object detection methods recognize objects…

计算机视觉与模式识别 · 计算机科学 2020-04-14 Qian Xie , Yu-Kun Lai , Jing Wu , Zhoutao Wang , Yiming Zhang , Kai Xu , Jun Wang

RGB-Thermal Salient Object Detection aims to pinpoint prominent objects within aligned pairs of visible and thermal infrared images. Traditional encoder-decoder architectures, while designed for cross-modality feature interactions, may not…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Hao Tang , Zechao Li , Dong Zhang , Shengfeng He , Jinhui Tang

Visual place classification from a first-person-view monocular RGB image is a fundamental problem in long-term robot navigation. A difficulty arises from the fact that RGB image classifiers are often vulnerable to spatial and appearance…

计算机视觉与模式识别 · 计算机科学 2023-05-12 Tomoya Iwasaki , Kanji Tanaka , Kenta Tsukahara

Real-time recognition of dynamic hand gestures from video streams is a challenging task since (i) there is no indication when a gesture starts and ends in the video, (ii) performed gestures should only be recognized once, and (iii) the…

计算机视觉与模式识别 · 计算机科学 2019-10-21 Okan Köpüklü , Ahmet Gunduz , Neslihan Kose , Gerhard Rigoll

Videos are inherently multimodal. This paper studies the problem of how to fully exploit the abundant multimodal clues for improved video categorization. We introduce a hybrid deep learning framework that integrates useful clues from…

多媒体 · 计算机科学 2017-06-15 Yu-Gang Jiang , Zuxuan Wu , Jinhui Tang , Zechao Li , Xiangyang Xue , Shih-Fu Chang

We look at the problem of developing a compact and accurate model for gesture recognition from videos in a deep-learning framework. Towards this we propose a joint 3DCNN-LSTM model that is end-to-end trainable and is shown to be better…

计算机视觉与模式识别 · 计算机科学 2018-01-01 Koustav Mullick , Anoop M. Namboodiri

Recent studies have demonstrated the power of recurrent neural networks for machine translation, image captioning and speech recognition. For the task of capturing temporal structure in video, however, there still remain numerous open…

计算机视觉与模式识别 · 计算机科学 2016-02-11 Lionel Pigou , Aäron van den Oord , Sander Dieleman , Mieke Van Herreweghe , Joni Dambre

In this paper, we propose a novel Attentive Multi-View Deep Subspace Nets (AMVDSN), which deeply explores underlying consistent and view-specific information from multiple views and fuse them by considering each view's dynamic contribution…

计算机视觉与模式识别 · 计算机科学 2021-12-24 Run-kun Lu , Jian-wei Liu , Xin Zuo

Scene recognition is one of the basic problems in computer vision research with extensive applications in robotics. When available, depth images provide helpful geometric cues that complement the RGB texture information and help to identify…

计算机视觉与模式识别 · 计算机科学 2021-09-08 Andrea Ferreri , Silvia Bucci , Tatiana Tommasi

Deep learning approaches have been established as the main methodology for video classification and recognition. Recently, 3-dimensional convolutions have been used to achieve state-of-the-art performance in many challenging video datasets.…

计算机视觉与模式识别 · 计算机科学 2020-06-24 Alexandros Stergiou , Georgios Kapidis , Grigorios Kalliatakis , Christos Chrysoulas , Remco Veltkamp , Ronald Poppe

Dense depth perception is critical for autonomous driving and other robotics applications. However, modern LiDAR sensors only provide sparse depth measurement. It is thus necessary to complete the sparse LiDAR data, where a synchronized…

计算机视觉与模式识别 · 计算机科学 2019-08-06 Jie Tang , Fei-Peng Tian , Wei Feng , Jian Li , Ping Tan

Visual scene understanding is an important capability that enables robots to purposefully act in their environment. In this paper, we propose a novel approach to object-class segmentation from multiple RGB-D views using deep learning. We…

计算机视觉与模式识别 · 计算机科学 2017-12-06 Lingni Ma , Jörg Stückler , Christian Kerl , Daniel Cremers

Compared to images, videos better reflect real-world acquisition and possess valuable temporal cues. However, existing multi-sensor fusion research predominantly integrates complementary context from multiple images rather than videos due…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Linfeng Tang , Yeda Wang , Meiqi Gong , Zizhuo Li , Yuxin Deng , Xunpeng Yi , Chunyu Li , Han Xu , Hao Zhang , Jiayi Ma

In this paper, we introduce a global video representation to video-based person re-identification (re-ID) that aggregates local 3D features across the entire video extent. Most of the existing methods rely on 2D convolutional networks…

计算机视觉与模式识别 · 计算机科学 2019-02-07 Lin Wu , Yang Wang , Ling Shao , Meng Wang