English
Related papers

Related papers: GEM: A Generalizable Ego-Vision Multimodal World M…

200 papers

Multimodal synthetic data generation is crucial in domains such as autonomous driving, robotics, augmented/virtual reality, and retail. We propose a novel approach, GenMM, for jointly editing RGB videos and LiDAR scans by inserting…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Bharat Singh , Viveka Kulharia , Luyu Yang , Avinash Ravichandran , Ambrish Tyagi , Ashish Shrivastava

Egocentric human videos provide scalable demonstrations for imitation learning, but existing corpora often lack either fine-grained, temporally localized action descriptions or dexterous hand annotations. We introduce OpenEgo, a multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Ahad Jawaid , Yu Xiang

Learned world models hold significant potential for robotic manipulation, as they can serve as simulator for real-world interactions. While extensive progress has been made in 2D video-based world models, these approaches often lack…

Robotics · Computer Science 2025-10-13 Chuanrui Zhang , Zhengxian Wu , Guanxing Lu , Yansong Tang , Ziwei Wang

End-to-end autonomous driving aims to generate safe and plausible planning policies from raw sensor input. Driving world models have shown great potential in learning rich representations by predicting the future evolution of a driving…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xingtai Gui , Meijie Zhang , Tianyi Yan , Wencheng Han , Jiahao Gong , Feiyang Tan , Cheng-zhong Xu , Jianbing Shen

Digital twin worlds with realistic interactive dynamics presents a new opportunity to develop generalist embodied agents in scannable environments with complex physical behaviors. To this end, we present GDGen (Generalized Representation…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Yichen Li , Zhiyi Li , Brandon Feng , Dinghuai Zhang , Antonio Torralba

World models serve as essential building blocks toward Artificial General Intelligence (AGI), enabling intelligent agents to predict future states and plan actions by simulating complex physical interactions. However, existing interactive…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Junyi Chen , Haoyi Zhu , Xianglong He , Yifan Wang , Jianjun Zhou , Wenzheng Chang , Yang Zhou , Zizun Li , Zhoujie Fu , Jiangmiao Pang , Tong He

Human motion prediction is important for many virtual and augmented reality (VR/AR) applications such as collision avoidance and realistic avatar generation. Existing methods have synthesised body motion only from observed past motion,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Haodong Yan , Zhiming Hu , Syn Schmitt , Andreas Bulling

Using an ego-centric camera to do localization and tracking is highly needed for urban navigation and indoor assistive system when GPS is not available or not accurate enough. The traditional hand-designed feature tracking and estimation…

Computer Vision and Pattern Recognition · Computer Science 2018-12-04 Liang Yang , Hao Jiang , Jizhong Xiao , Zhouyuan Huo

In many applications, the training data for a machine learning task is partitioned across multiple nodes, and aggregating this data may be infeasible due to communication, privacy, or storage constraints. Existing distributed optimization…

Machine Learning · Computer Science 2019-06-06 Neel Guha , Virginia Smith

Multi-modal end-to-end autonomous driving has shown promising advancements in recent work. By embedding more modalities into end-to-end networks, the system's understanding of both static and dynamic aspects of the driving environment is…

Robotics · Computer Science 2025-05-15 Ziang Guo , Xinhao Lin , Zakhar Yagudin , Artem Lykov , Yong Wang , Yanqiang Li , Dzmitry Tsetserukou

High-quality driving video generation is crucial for providing training data for autonomous driving models. However, current generative models rarely focus on enhancing camera motion control under multi-view tasks, which is essential for…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Yining Yao , Xi Guo , Chenjing Ding , Wei Wu

Humans build viewpoint-independent cognitive maps through navigation, enabling intuitive reasoning about object permanence and spatial relations. We argue that multimodal large language models (MLLMs), despite extensive video training, lack…

Machine Learning · Computer Science 2025-12-02 Jacob Thompson , Emiliano Garcia-Lopez , Yonatan Bisk

Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Hila Chefer , Uriel Singer , Amit Zohar , Yuval Kirstain , Adam Polyak , Yaniv Taigman , Lior Wolf , Shelly Sheynin

The integration of brain-computer interfaces (BCIs), in particular electroencephalography (EEG), with artificial intelligence (AI) has shown tremendous promise in decoding human cognition and behavior from neural signals. In particular, the…

Artificial Intelligence · Computer Science 2025-10-15 Nie Lin , Yansen Wang , Dongqi Han , Weibang Jiang , Jingyuan Li , Ryosuke Furuta , Yoichi Sato , Dongsheng Li

End-to-end autonomous driving planners typically generate trajectories from current observations alone. However, real-world driving is highly dynamic, and such reactive planning cannot anticipate future scene evolution, often leading to…

Robotics · Computer Science 2026-04-29 Chuyao Fu , Shengzhe Gan , Zhuoli Ouyang , Yuhan Rui , Xiaowei Chi , Sirui Han , Jiankun Wang , Hong Zhang

Multi-view egocentric dynamic scene reconstruction holds significant research value for applications in holographic documentation of social interactions. However, existing reconstruction datasets focus on static multi-view or…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Bate Li , Houqiang Zhong , Zhengxue Cheng , Qiang Hu , Qiang Wang , Li Song , Wenjun Zhang

Egocentric vision is essential for both human and machine visual understanding, particularly in capturing the detailed hand-object interactions needed for manipulation tasks. Translating third-person views into first-person views…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Junho Park , Andrew Sangwoo Ye , Taein Kwon

Predicting the future to anticipate the outcome of events and actions is a critical attribute of autonomous agents; particularly for agents which must rely heavily on real time visual data for decision making. Working towards this…

Computer Vision and Pattern Recognition · Computer Science 2018-11-29 Suhani Vora , Reza Mahjourian , Soeren Pirk , Anelia Angelova

Understanding social interactions from egocentric views is crucial for many applications, ranging from assistive robotics to AR/VR. Key to reasoning about interactions is to understand the body pose and motion of the interaction partner…

Computer Vision and Pattern Recognition · Computer Science 2022-08-17 Siwei Zhang , Qianli Ma , Yan Zhang , Zhiyin Qian , Taein Kwon , Marc Pollefeys , Federica Bogo , Siyu Tang

We present SpatialMem, a memory-centric system for long-horizon, language-grounded retrieval and QA from egocentric video, where metric 3D serves as an interpretable indexing scaffold rather than an explicit mapping objective. Starting from…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Xinyi Zheng , Yunze Liu , Chi-Hao Wu , Fan Zhang , Hao Zheng , Wenqi Zhou , Walterio W. Mayol-Cuevas , Junxiao Shen