English
Related papers

Related papers: Jointly Modeling Motion and Appearance Cues for Ro…

200 papers

Joint Detection and Embedding (JDE) trackers have demonstrated excellent performance in Multi-Object Tracking (MOT) tasks by incorporating the extraction of appearance features as auxiliary tasks through embedding Re-Identification task…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Yunfei Zhang , Chao Liang , Jin Gao , Zhipeng Zhang , Weiming Hu , Stephen Maybank , Xue Zhou , Liang Li

2D face recognition encounters challenges in unconstrained environments due to varying illumination, occlusion, and pose. Recent studies focus on RGB-D face recognition to improve robustness by incorporating depth information. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Zijian Chen , Mei Wang , Weihong Deng , Hongzhi Shi , Dongchao Wen , Yingjie Zhang , Xingchen Cui , Jian Zhao

Efficiently exploiting multi-modal inputs for accurate RGB-D saliency detection is a topic of high interest. Most existing works leverage cross-modal interactions to fuse the two streams of RGB-D for intermediate features' enhancement. In…

Computer Vision and Pattern Recognition · Computer Science 2022-08-31 Zongwei Wu , Shriarulmozhivarman Gobichettipalayam , Brahim Tamadazte , Guillaume Allibert , Danda Pani Paudel , Cédric Demonceaux

Multiple object tracking (MOT) is a significant task in achieving autonomous driving. Traditional works attempt to complete this task, either based on point clouds (PC) collected by LiDAR, or based on images captured from cameras. However,…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Guangming Wang , Chensheng Peng , Jinpeng Zhang , Hesheng Wang

We present a method for temporally consistent motion segmentation from RGB-D videos assuming a piecewise rigid motion model. We formulate global energies over entire RGB-D sequences in terms of the segmentation of each frame into a number…

Computer Vision and Pattern Recognition · Computer Science 2016-08-17 Peter Bertholet , Alexandru-Eugen Ichim , Matthias Zwicker

Recently, many multi-modal trackers prioritize RGB as the dominant modality, treating other modalities as auxiliary, and fine-tuning separately various multi-modal tasks. This imbalance in modality dependence limits the ability of methods…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Xiantao Hu , Bineng Zhong , Qihua Liang , Zhiyi Mo , Liangtao Shi , Ying Tai , Jian Yang

In the paper, we propose a robust real-time visual odometry in dynamic environments via rigid-motion model updated by scene flow. The proposed algorithm consists of spatial motion segmentation and temporal motion tracking. The spatial…

Robotics · Computer Science 2019-07-22 Sangil Lee , Clark Youngdong Son , H. Jin Kim

The integration of dual-modal features has been pivotal in advancing RGB-Depth (RGB-D) tracking. However, current trackers are less efficient and focus solely on single-level features, resulting in weaker robustness in fusion and slower…

Computer Vision and Pattern Recognition · Computer Science 2025-04-25 Boyue Xu , Yi Xu , Ruichao Hou , Jia Bei , Tongwei Ren , Gangshan Wu

Although fusing multiple sensor modalities can enhance object detection performance, existing fusion approaches often overlook subtle variations in environmental conditions and sensor inputs. As a result, they struggle to adaptively weight…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Aditya Taparia , Noel Ngu , Mario Leiva , Joshua Shay Kricheli , John Corcoran , Nathaniel D. Bastian , Gerardo Simari , Paulo Shakarian , Ransalu Senanayake

RGB-D saliency detection integrates information from both RGB images and depth maps to improve prediction of salient regions under challenging conditions. The key to RGB-D saliency detection is to fully mine and fuse information at multiple…

Computer Vision and Pattern Recognition · Computer Science 2021-12-02 Yue Wang , Xu Jia , Lu Zhang , Yuke Li , James Elder , Huchuan Lu

We introduce a novel robust hybrid 3D face tracking framework from RGBD video streams, which is capable of tracking head pose and facial actions without pre-calibration or intervention from a user. In particular, we emphasize on improving…

Computer Vision and Pattern Recognition · Computer Science 2015-07-13 Hai X. Pham , Chongyu Chen , Luc N. Dao , Vladimir Pavlovic , Jianfei Cai , Tat-jen Cham

A robust algorithm solution is proposed for tracking an object in complex video scenes. In this solution, the bootstrap particle filter (PF) is initialized by an object detector, which models the time-evolving background of the video signal…

Computer Vision and Pattern Recognition · Computer Science 2015-09-29 Yi Dai , Bin Liu

RGB-D salient object detection aims to identify the most visually distinctive objects in a pair of color and depth images. Based upon an observation that most of the salient objects may stand out at least in one modality, this paper…

Computer Vision and Pattern Recognition · Computer Science 2019-01-09 Ningning Wang , Xiaojin Gong

How to make a good trade-off between performance and computational cost is crucial for a tracker. However, current famous methods typically focus on complicated and time-consuming learning that combining temporal and appearance information…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Jinxia Xie , Bineng Zhong , Qihua Liang , Ning Li , Zhiyi Mo , Shuxiang Song

RGBD object tracking is gaining momentum in computer vision research thanks to the development of depth sensors. Although numerous RGBD trackers have been proposed with promising performance, an in-depth review for comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Jinyu Yang , Zhe Li , Song Yan , Feng Zheng , Aleš Leonardis , Joni-Kristian Kämäräinen , Ling Shao

In multi-modal action recognition, it is important to consider not only the complementary nature of different modalities but also global action content. In this paper, we propose a novel network, named Modality Mixer (M-Mixer) network, to…

Computer Vision and Pattern Recognition · Computer Science 2023-02-22 Sumin Lee , Sangmin Woo , Yeonju Park , Muhammad Adi Nugroho , Changick Kim

We address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from most existing methods that only consider RGB images as…

Computer Vision and Pattern Recognition · Computer Science 2021-11-01 Yi-Wen Chen , Yi-Hsuan Tsai , Ming-Hsuan Yang

A significant challenge in object detection is accurate identification of an object's position in image space, whereas one algorithm with one set of parameters is usually not enough, and the fusion of multiple algorithms and/or parameters…

Computer Vision and Pattern Recognition · Computer Science 2018-03-20 Pan Wei , John E. Ball , Derek T. Anderson

The emergence of different sensors (Near-Infrared, Depth, etc.) is a remedy for the limited application scenarios of traditional RGB camera. The RGB-X tasks, which rely on RGB input and another type of data input to resolve specific…

Computer Vision and Pattern Recognition · Computer Science 2023-06-23 Jin Ma , Jinlong Li , Qing Guo , Tianyun Zhang , Yuewei Lin , Hongkai Yu

The research introduces a reproducible framework for transforming raw, heterogeneous sensor streams into aligned, semantically meaningful representations for multimodal human activity recognition. Grounded in the Carnegie Mellon University…

Applications · Statistics 2026-05-05 Yiyao Yang , Yasemin Gulbahar