English
Related papers

Related papers: GEM: A Generalizable Ego-Vision Multimodal World M…

200 papers

Models for egocentric 3D and 4D reconstruction, including few-shot interpolation and extrapolation settings, can benefit from having images from exocentric viewpoints as supervision signals. No existing dataset provides the necessary…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 Marius Kästingschäfer , Théo Gieruc , Sebastian Bernhard , Dylan Campbell , Eldar Insafutdinov , Eyvaz Najafli , Thomas Brox

We introduce Latent-WAM, an efficient end-to-end autonomous driving framework that achieves strong trajectory planning through spatially-aware and dynamics-informed latent world representations. Existing world-model-based planners suffer…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Linbo Wang , Yupeng Zheng , Qiang Chen , Shiwei Li , Yichen Zhang , Zebin Xing , Qichao Zhang , Xiang Li , Deheng Qian , Pengxuan Yang , Yihang Dong , Ce Hao , Xiaoqing Ye , Junyu han , Yifeng Pan , Dongbin Zhao

Generalist robot policies can now perform a wide range of manipulation skills, but evaluating and improving their ability with unfamiliar objects and instructions remains a significant challenge. Rigorous evaluation requires a large number…

Robotics · Computer Science 2026-03-03 Yanjiang Guo , Lucy Xiaoyang Shi , Jianyu Chen , Chelsea Finn

Multimodal foundation models that can holistically process text alongside images, video, audio, and other sensory modalities are increasingly used in a variety of real-world applications. However, it is challenging to characterize and study…

Smart glass is emerging as an useful device since it provides plenty of insights under hands-busy, eyes-on-task situations. To understand the context of the wearer, 6D object pose estimation in egocentric view is becoming essential.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Taegyoon Yoon , Yegyu Han , Seojin Ji , Jaewoo Park , Sojeong Kim , Taein Kwon , Hyung-Sin Kim

Video depth estimation extends monocular prediction into the temporal domain to ensure coherence. However, existing methods often suffer from spatial blurring in fine-detail regions and temporal inconsistencies. We argue that current…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Yuecheng Liu , Junda Cheng , Longliang Liu , Wenjing Liao , Hanrui Cheng , Yuzhou Wang , Xin Yang

World models have made significant progress in modeling dynamic environments; however, most embodied world models are still restricted to 2D representations, lacking the comprehensive multi-view information essential for embodied spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Peiyan Tu , Hanxin Zhu , Jingwen Sun , Shaojie Ren , Cong Wang , Jiayi Luo , Xiaoqian Cheng , Zhibo Chen

Learning-based, single-view depth estimation often generalizes poorly to unseen datasets. While learning-based, two-frame depth estimation solves this problem to some extent by learning to match features across frames, it performs poorly at…

Computer Vision and Pattern Recognition · Computer Science 2018-05-18 Rui Wang , Jan-Michael Frahm , Stephen M. Pizer

Tabular data dominates data science but poses challenges for generative models, especially when the data is limited or sensitive. We present a novel approach to generating synthetic tabular data based on the principle of maximum entropy --…

Machine Learning · Computer Science 2025-09-23 Miao Li , Phuc Nguyen , Christopher Tam , Alexandra Morgan , Kenneth Ge , Rahul Bansal , Linzi Yu , Rima Arnaout , Ramy Arnaout

End-to-end autonomous driving models increasingly benefit from large vision--language models for semantic understanding, yet ensuring safe and accurate operation under long-tail conditions remains challenging. These challenges are…

Robotics · Computer Science 2026-02-03 Weizhe Tang , Junwei You , Jiaxi Liu , Zhaoyi Wang , Rui Gan , Zilin Huang , Feng Wei , Bin Ran

Recently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical applications and failing…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Chao Xu , Junwei Zhu , Jiangning Zhang , Yue Han , Wenqing Chu , Ying Tai , Chengjie Wang , Zhifeng Xie , Yong Liu

Full-body egocentric pose estimation from head and hand poses alone has become an active area of research to power articulate avatar representations on headset-based platforms. However, existing methods over-rely on the indoor…

Computer Vision and Pattern Recognition · Computer Science 2024-09-09 Jiaxi Jiang , Paul Streli , Manuel Meier , Christian Holz

Future trajectory prediction of a tracked pedestrian from an egocentric perspective is a key task in areas such as autonomous driving and robot navigation. The challenge of this task lies in the complex dynamic relative motion between the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yusheng Peng , Gaofeng Zhang , Liping Zheng

A self-driving perception model aims to extract 3D semantic representations from multiple cameras collectively into the bird's-eye-view (BEV) coordinate frame of the ego car in order to ground downstream planner. Existing perception methods…

Computer Vision and Pattern Recognition · Computer Science 2022-08-19 Jiachen Lu , Zheyuan Zhou , Xiatian Zhu , Hang Xu , Li Zhang

Learning an agent model that behaves like humans-capable of jointly perceiving the environment, predicting the future, and taking actions from a first-person perspective-is a fundamental challenge in computer vision. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Lu Chen , Yizhou Wang , Shixiang Tang , Qianhong Ma , Tong He , Wanli Ouyang , Xiaowei Zhou , Hujun Bao , Sida Peng

Recent advances in 3D foundation models have led to growing interest in reconstructing humans and their surrounding environments. However, most existing approaches focus on monocular inputs, and extending them to multi-view settings…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Sangmin Kim , Minhyuk Hwang , Geonho Cha , Dongyoon Wee , Jaesik Park

Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Chengjun Yu , Xuhan Zhu , Chaoqun Du , Pengfei Yu , Wei Zhai , Yang Cao , Zheng-Jun Zha

Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made them powerful tools in embodied navigation, enabling agents to leverage commonsense and spatial reasoning for efficient exploration in…

Egocentric 3D hand pose estimation and gesture recognition are essential for immersive augmented/virtual reality, human-computer interaction, and robotics. However, conventional frame-based cameras suffer from motion blur and limited…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Luming Wang , Hao Shi , Jiajun Zhai , Kailun Yang , Kaiwei Wang

The fusion of sensor data from heterogeneous sensors is crucial for robust perception in various robotics applications that involve moving platforms, for instance, autonomous vehicle navigation. In particular, combining camera and lidar…

Robotics · Computer Science 2020-03-10 Mao Shan , Julie Stephany Berrio , Stewart Worrall , Eduardo Nebot