English
Related papers

Related papers: GHOST: Ground-projected Hypotheses from Observed S…

200 papers

Object tracking is central to robot perception and scene understanding. Tracking-by-detection has long been a dominant paradigm for object tracking of specific object categories. Recently, large-scale pre-trained models have shown promising…

Computer Vision and Pattern Recognition · Computer Science 2024-01-26 Wen-Hsuan Chu , Adam W. Harley , Pavel Tokmakov , Achal Dave , Leonidas Guibas , Katerina Fragkiadaki

We consider the problem of vision-based pose estimation for autonomous systems. While deep neural networks have been successfully used for vision-based tasks, they inherently lack provable guarantees on the correctness of their output,…

Robotics · Computer Science 2026-01-27 Ulices Santa Cruz , Mahmoud Elfar , Yasser Shoukry

Planning the trajectory of the controlled ego vehicle is a key challenge in automated driving. As for human drivers, predicting the motions of surrounding vehicles is important to plan the own actions. Recent motion prediction methods…

Robotics · Computer Science 2024-03-19 Steffen Hagedorn , Marcel Milich , Alexandru P. Condurache

Accurate perception of the vehicle's 3D surroundings, including fine-scale road geometry, such as bumps, slopes, and surface irregularities, is essential for safe and comfortable vehicle control. However, conventional monocular depth…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Gasser Elazab , Maximilian Jansen , Michael Unterreiner , Olaf Hellwich

Accurate observation of dynamic environments traditionally relies on synthesizing raw, signal-level information from multiple distributed sensors. This work investigates an alternative approach: performing geospatial inference using only…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Sadik Yagiz Yetim , Gaofeng Dong , Isaac-Neil Zanoria , Ronit Barman , Maggie Wigness , Tarek Abdelzaher , Mani Srivastava , Suhas Diggavi

We propose Video Gaussian Masked Autoencoders (Video-GMAE), a self-supervised approach for representation learning that encodes a sequence of images into a set of Gaussian splats moving over time. Representing a video as a set of Gaussians…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Tanish Baranwal , Himanshu Gaurav Singh , Jathushan Rajasegaran , Jitendra Malik

We propose an algorithm for real-time 6DOF pose tracking of rigid 3D objects using a monocular RGB camera. The key idea is to derive a region-based cost function using temporally consistent local color histograms. While such region-based…

Computer Vision and Pattern Recognition · Computer Science 2018-12-20 Henning Tjaden , Ulrich Schwanecke , Elmar Schömer , Daniel Cremers

Although considerable advancements have been attained in self-supervised depth estimation from monocular videos, most existing methods often treat all objects in a video as static entities, which however violates the dynamic nature of…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Xiuzhe Wu , Xiaoyang Lyu , Qihao Huang , Yong Liu , Yang Wu , Ying Shan , Xiaojuan Qi

Modeling and rendering dynamic urban driving scenes is crucial for self-driving simulation. Current high-quality methods typically rely on costly manual object tracklet annotations, while self-supervised approaches fail to capture dynamic…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Jiawei Xu , Kai Deng , Zexin Fan , Shenlong Wang , Jin Xie , Jian Yang

We propose a stereo vision-based approach for tracking the camera ego-motion and 3D semantic objects in dynamic autonomous driving scenarios. Instead of directly regressing the 3D bounding box using end-to-end approaches, we propose to use…

Computer Vision and Pattern Recognition · Computer Science 2018-11-30 Peiliang Li , Tong Qin , Shaojie Shen

For ego-motion estimation, the feature representation of the scenes is crucial. Previous methods indicate that both the low-level and semantic feature-based methods can achieve promising results. Therefore, the incorporation of hierarchical…

Computer Vision and Pattern Recognition · Computer Science 2019-08-06 Xiaochuan Yin , Chengju Liu

We propose an end-to-end learning framework for segmenting generic objects in both images and videos. Given a novel image or video, our approach produces a pixel-level mask for all "object-like" regions---even for object categories never…

Computer Vision and Pattern Recognition · Computer Science 2018-12-19 Bo Xiong , Suyog Dutt Jain , Kristen Grauman

We introduce a novel framework to build a model that can learn how to segment objects from a collection of images without any human annotation. Our method builds on the observation that the location of object segments can be perturbed…

Computer Vision and Pattern Recognition · Computer Science 2019-11-05 Adam Bielski , Paolo Favaro

Fisheye cameras are commonly used in applications like autonomous driving and surveillance to provide a large field of view ($>180^{\circ}$). However, they come at the cost of strong non-linear distortions which require more complex…

Computer Vision and Pattern Recognition · Computer Science 2020-10-08 Varun Ravi Kumar , Sandesh Athni Hiremath , Stefan Milz , Christian Witt , Clement Pinnard , Senthil Yogamani , Patrick Mader

In order to track the moving objects in long range against occlusion, interruption, and background clutter, this paper proposes a unified approach for global trajectory analysis. Instead of the traditional frame-by-frame tracking, our…

Computer Vision and Pattern Recognition · Computer Science 2015-02-03 Liang Lin , Yongyi Lu , Yan Pan , Xiaowu Chen

Prior point cloud provides 3D environmental context, which enhances the capabilities of monocular camera in downstream vision tasks, such as 3D object detection, via data fusion. However, the absence of accurate and automated registration…

Robotics · Computer Science 2024-04-09 Yu Sheng , Lu Zhang , Xingchen Li , Yifan Duan , Yanyong Zhang , Yu Zhang , Jianmin Ji

We introduce Vid-CamEdit, a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Junyoung Seo , Jisang Han , Jaewoo Jung , Siyoon Jin , Joungbin Lee , Takuya Narihira , Kazumi Fukuda , Takashi Shibuya , Donghoon Ahn , Shoukang Hu , Seungryong Kim , Yuki Mitsufuji

This paper presents a novel yet intuitive approach to unsupervised feature learning. Inspired by the human visual system, we explore whether low-level motion-based grouping cues can be used to learn an effective visual representation.…

Computer Vision and Pattern Recognition · Computer Science 2017-04-13 Deepak Pathak , Ross Girshick , Piotr Dollár , Trevor Darrell , Bharath Hariharan

Unsupervised learning of depth and ego-motion from unlabelled monocular videos has recently drawn great attention, which avoids the use of expensive ground truth in the supervised one. It achieves this by using the photometric errors…

Computer Vision and Pattern Recognition · Computer Science 2020-11-23 Hualie Jiang , Laiyan Ding , Zhenglong Sun , Rui Huang

We introduce the Visual Implicit Geometry Transformer (ViGT), an autonomous driving geometric model that estimates continuous 3D occupancy fields from surround-view camera rigs. ViGT represents a step towards foundational geometric models…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Arsenii Shirokov , Mikhail Kuznetsov , Danila Stepochkin , Egor Evdokimov , Daniil Glazkov , Nikolay Patakin , Anton Konushin , Dmitry Senushkin
‹ Prev 1 3 4 5 6 7 10 Next ›