English
Related papers

Related papers: MonoMobility: Zero-Shot 3D Mobility Analysis from …

200 papers

Photorealistic reconstruction of street scenes is essential for developing real-world simulators in autonomous driving. While recent methods based on 3D/4D Gaussian Splatting (GS) have demonstrated promising results, they still encounter…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Xiaobao Wei , Qingpo Wuwu , Zhongyu Zhao , Zhuangzhe Wu , Nan Huang , Ming Lu , Ningning MA , Shanghang Zhang

We propose a novel approach for unsupervised 3D animation of non-rigid deformable objects. Our method learns the 3D structure and dynamics of objects solely from single-view RGB videos, and can decompose them into semantically meaningful…

Computer Vision and Pattern Recognition · Computer Science 2023-01-27 Aliaksandr Siarohin , Willi Menapace , Ivan Skorokhodov , Kyle Olszewski , Jian Ren , Hsin-Ying Lee , Menglei Chai , Sergey Tulyakov

In this paper, we tackle the problem of multibody SLAM from a monocular camera. The term multibody, implies that we track the motion of the camera, as well as that of other dynamic participants in the scene. The quintessential challenge in…

Future robots are envisioned as versatile systems capable of performing a variety of household tasks. The big question remains, how can we bridge the embodiment gap while minimizing physical robot learning, which fundamentally does not…

Robotics · Computer Science 2025-03-31 Hanzhi Chen , Boyang Sun , Anran Zhang , Marc Pollefeys , Stefan Leutenegger

Zero-shot action recognition, which recognizes actions in videos without having received any training examples, is gaining wide attention considering it can save labor costs and training time. Nevertheless, the performance of zero-shot…

Computer Vision and Pattern Recognition · Computer Science 2023-06-13 Nan Wu , Hiroshi Kera , Kazuhiko Kawamoto

Enabling robots to execute novel manipulation tasks zero-shot is a central goal in robotics. Most existing methods assume in-distribution tasks or rely on fine-tuning with embodiment-matched data, limiting transfer across platforms. We…

Robotics · Computer Science 2025-10-10 Hongyu Li , Lingfeng Sun , Yafei Hu , Duy Ta , Jennifer Barry , George Konidaris , Jiahui Fu

Recent advances in 4D scene reconstruction have significantly improved dynamic modeling across various domains. However, existing approaches remain limited under aerial conditions with single-view capture, wide spatial range, and dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Hanyang Liu , Rongjun Qin

It has long been challenging to recover the underlying dynamic 3D scene representations from a monocular RGB video. Existing works formulate this problem into finding a single most plausible solution by adding various constraints such as…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Ziyang Song , Jinxi Li , Bo Yang

We present a system that allows for accurate, fast, and robust estimation of camera parameters and depth maps from casual monocular videos of dynamic scenes. Most conventional structure from motion and monocular SLAM techniques assume input…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Zhengqi Li , Richard Tucker , Forrester Cole , Qianqian Wang , Linyi Jin , Vickie Ye , Angjoo Kanazawa , Aleksander Holynski , Noah Snavely

Learning to estimate 3D geometry in a single image by watching unlabeled videos via deep convolutional network has made significant process recently. Current state-of-the-art (SOTA) methods, are based on the learning framework of rigid…

Computer Vision and Pattern Recognition · Computer Science 2018-08-17 Zhenheng Yang , Peng Wang , Yang Wang , Wei Xu , Ram Nevatia

Reconstructing people, objects, and their interactions in 3D is a long-standing goal for intelligent systems. Often the input is RGB video from a moving camera, making the task ill-posed; depth is ambiguous, humans and objects occlude each…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Lixin Xue , Chengwei Zheng , Georgios Paschalidis , Chen Guo , Manuel Kaufmann , Juan Zarate , Dimitrios Tzionas

We propose a stereo vision-based approach for tracking the camera ego-motion and 3D semantic objects in dynamic autonomous driving scenarios. Instead of directly regressing the 3D bounding box using end-to-end approaches, we propose to use…

Computer Vision and Pattern Recognition · Computer Science 2018-11-30 Peiliang Li , Tong Qin , Shaojie Shen

Current motion-based multiple object tracking (MOT) approaches rely heavily on Intersection-over-Union (IoU) for object association. Without using 3D features, they are ineffective in scenarios with occlusions or visually similar objects.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Milad Khanchi , Maria Amer , Charalambos Poullis

Multi-modal 3D object understanding has gained significant attention, yet current approaches often assume complete data availability and rigid alignment across all modalities. We present CrossOver, a novel framework for cross-modal 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Sayan Deb Sarkar , Ondrej Miksik , Marc Pollefeys , Daniel Barath , Iro Armeni

Despite advancements in Multimodal Large Language Models (MLLMs), their proficiency in fine-grained video motion understanding remains critically limited. They often lack inter-frame differencing and tend to average or ignore subtle visual…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Yipeng Du , Tiehan Fan , Kepan Nan , Rui Xie , Penghao Zhou , Xiang Li , Jian Yang , Zhenheng Yang , Ying Tai

We propose Track and Caption Any Motion (TCAM), a motion-centric framework for automatic video understanding that discovers and describes motion patterns without user queries. Understanding videos in challenging conditions like occlusion,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Bishoy Galoaa , Sarah Ostadabbas

Dynamic Object-aware SLAM (DOS) exploits object-level information to enable robust motion estimation in dynamic environments. Existing methods mainly focus on identifying and excluding dynamic objects from the optimization. In this paper,…

Robotics · Computer Science 2022-11-15 Yuheng Qiu , Chen Wang , Wenshan Wang , Mina Henein , Sebastian Scherer

For ego-motion estimation, the feature representation of the scenes is crucial. Previous methods indicate that both the low-level and semantic feature-based methods can achieve promising results. Therefore, the incorporation of hierarchical…

Computer Vision and Pattern Recognition · Computer Science 2019-08-06 Xiaochuan Yin , Chengju Liu

We propose a learning-based method that solves monocular stereo and can be extended to fuse depth information from multiple target frames. Given two unconstrained images from a monocular camera with known intrinsic calibration, our network…

Computer Vision and Pattern Recognition · Computer Science 2019-09-13 Kaixuan Wang , Shaojie Shen

We demonstrate that, under orthographic projection and with a camera fixated on a point located on a rigid body, the rotation of that body can be analytically obtained by tracking only one other feature in the image. With some exceptions,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Daniel Raviv , Juan D. Yepes , Eiki M. Martinson
‹ Prev 1 8 9 10 Next ›