English
Related papers

Related papers: Geometry-Aware Rotary Position Embedding for Consi…

200 papers

Visual repetition is ubiquitous in our world. It appears in human activity (sports, cooking), animal behavior (a bee's waggle dance), natural phenomena (leaves in the wind) and in urban environments (flashing lights). Estimating visual…

Computer Vision and Pattern Recognition · Computer Science 2018-06-20 Tom F. H. Runia , Cees G. M. Snoek , Arnold W. M. Smeulders

Understanding and reconstructing the complex geometry and motion of dynamic scenes from video remains a formidable challenge in computer vision. This paper introduces D4RT, a simple yet powerful feedforward model designed to efficiently…

Estimating the 6D pose of textureless objects from RGB images is an important problem in robotics. Due to appearance ambiguities, rotational symmetries, and severe occlusions, single-view based 6D pose estimators are still unable to handle…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Jun Yang , Wenjie Xue , Sahar Ghavidel , Steven L. Waslander

Incorporating inductive bias by embedding geometric entities (such as rays) as input has proven successful in multi-view learning. However, the methods adopting this technique typically lack equivariance, which is crucial for effective 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Yinshuang Xu , Dian Chen , Katherine Liu , Sergey Zakharov , Rares Ambrus , Kostas Daniilidis , Vitor Guizilini

Modeling 4D scenes requires capturing both spatial structure and temporal motion, which is challenging due to the need for physically consistent representations of complex rigid and non-rigid motions. Existing approaches mainly rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Weidong Qiao , Wangmeng Zuo , Hui Li

Vision-Language-Action (VLA) models often fail to generalize to unseen camera viewpoints, a limitation stemming from their difficulty in inferring robust 3D geometry from 2D images. We introduce GeoAware-VLA, a simple yet effective approach…

Robotics · Computer Science 2026-03-10 Ali Abouzeid , Malak Mansour , Qinbo Sun , Zezhou Sun , Dezhen Song

Understanding the 3-dimensional structure of the world is a core challenge in computer vision and robotics. Neural rendering approaches learn an implicit 3D model by predicting what a camera would see from an arbitrary viewpoint. We extend…

Computer Vision and Pattern Recognition · Computer Science 2019-11-13 Josh Tobin , OpenAI Robotics , Pieter Abbeel

Purpose: Surgical scene understanding plays a critical role in the technology stack of tomorrow's intervention-assisting systems in endoscopic surgeries. For this, tracking the endoscope pose is a key component, but remains challenging due…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Michel Hayoz , Christopher Hahne , Mathias Gallardo , Daniel Candinas , Thomas Kurmann , Maximilian Allan , Raphael Sznitman

Inferring geometrically consistent dense 3D scenes across a tuple of temporally consecutive images remains challenging for self-supervised monocular depth prediction pipelines. This paper explores how the increasingly popular transformer…

Computer Vision and Pattern Recognition · Computer Science 2021-10-18 Patrick Ruhkamp , Daoyi Gao , Hanzhi Chen , Nassir Navab , Benjamin Busam

Camera pose estimation in known scenes is a 3D geometry task recently tackled by multiple learning algorithms. Many regress precise geometric quantities, like poses or 3D points, from an input image. This either fails to generalize to new…

Automated 3D scene generation is pivotal for applications spanning virtual reality, digital content creation, and Embodied AI. While computer graphics prioritizes aesthetic layouts, vision and robotics demand scenes that mirror real-world…

Graphics · Computer Science 2026-03-31 Minzhang Li , Kuixiang Shao , Xuebing Li , Yuyang Jiao , Yinuo Bai , Hengan Zhou , Sixian Shen , Jiayuan Gu , Jingyi Yu

Photo-realistic rendering and novel view synthesis play a crucial role in human-computer interaction tasks, from gaming to path planning. Neural Radiance Fields (NeRFs) model scenes as continuous volumetric functions and achieve remarkable…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Iryna Repinetska , Anna Hilsmann , Peter Eisert

We propose GazeNeRF, a 3D-aware method for the task of gaze redirection. Existing gaze redirection methods operate on 2D images and struggle to generate 3D consistent results. Instead, we build on the intuition that the face region and…

Computer Vision and Pattern Recognition · Computer Science 2023-03-30 Alessandro Ruzzi , Xiangwei Shi , Xi Wang , Gengyan Li , Shalini De Mello , Hyung Jin Chang , Xucong Zhang , Otmar Hilliges

This paper proposes AutoScape, a long-horizon driving scene generation framework. At its core is a novel RGB-D diffusion model that iteratively generates sparse, geometrically consistent keyframes, serving as reliable anchors for the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Jiacheng Chen , Ziyu Jiang , Mingfu Liang , Bingbing Zhuang , Jong-Chyi Su , Sparsh Garg , Ying Wu , Manmohan Chandraker

We propose a novel 3D gaze estimation approach that learns spatial relationships between the subject and objects in the scene, and outputs 3D gaze direction. Our method targets unconstrained settings, including cases where close-up views of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Yuki Kawana , Shintaro Shiba , Quan Kong , Norimasa Kobori

Reconstructing dynamic 3D scenes from sparse multi-view videos is highly ill-posed, often leading to geometric collapse, trajectory drift, and floating artifacts. Recent attempts introduce generative priors to hallucinate missing content,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Zhenlong Wu , Zihan Zheng , Xuanxuan Wang , Qianhe Wang , Hua Yang , Xiaoyun Zhang , Qiang Hu , Wenjun Zhang

Face animation aims at creating photo-realistic portrait videos with animated poses and expressions. A common practice is to generate displacement fields that are used to warp pixels and features from source to target. However, prior…

Computer Vision and Pattern Recognition · Computer Science 2022-10-06 Yatao Zhong , Faezeh Amjadi , Ilya Zharkov

Precise geolocalization is crucial for unmanned aerial vehicles (UAVs). However, most current deployed UAVs rely on the global navigation satellite systems (GNSS) or high precision inertial navigation systems (INS) for geolocalization. In…

Robotics · Computer Science 2023-01-02 Jun Mao , Lilian Zhang , Xiaofeng He , Hao Qu , Xiaoping Hu

Recent studies have shown remarkable advances in 3D human pose estimation from monocular images, with the help of large-scale in-door 3D datasets and sophisticated network architectures. However, the generalizability to different…

Computer Vision and Pattern Recognition · Computer Science 2019-03-28 Xipeng Chen , Kwan-Yee Lin , Wentao Liu , Chen Qian , Xiaogang Wang , Liang Lin

Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rather than passively reconstructing scenes from pre-collected images. However, existing active…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Tianling Xu , Shengzhe Gan , Leslie Gu , Yuelei Li , Fangneng Zhan , Hanspeter Pfister
‹ Prev 1 8 9 10 Next ›