English
Related papers

Related papers: GEM: A Generalizable Ego-Vision Multimodal World M…

200 papers

Predicting future human behavior from an input human video is a useful task for applications such as autonomous driving and robotics. While most previous works predict a single future, multiple futures with different behavior can…

Computer Vision and Pattern Recognition · Computer Science 2021-06-02 Naoya Fushishita , Antonio Tejero-de-Pablos , Yusuke Mukuta , Tatsuya Harada

Recent Multimodal Large Language Models (MLLMs) have shown high potential for spatial reasoning within 3D scenes. However, they typically rely on computationally expensive 3D representations like point clouds or reconstructed Bird's-Eye…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Shuyao Shi , Kang G. Shin

We present an approach which takes advantage of both structure and semantics for unsupervised monocular learning of depth and ego-motion. More specifically, we model the motion of individual objects and learn their 3D motion vector jointly…

Computer Vision and Pattern Recognition · Computer Science 2019-06-14 Vincent Casser , Soeren Pirk , Reza Mahjourian , Anelia Angelova

In multi-vector retrieval, both queries and data are represented as sets of high-dimensional vectors, enabling finer-grained semantic matching and improving retrieval quality over single-vector approaches. However, its practical adoption is…

Information Retrieval · Computer Science 2026-03-24 Yao Tian , Zhoujin Tian , Xi Zhao , Ruiyuan Zhang , Xiaofang Zhou

Training self-driving systems to be robust to the long-tail of driving scenarios is a critical problem. Model-based approaches leverage simulation to emulate a wide range of scenarios without putting users at risk in the real world. One…

Robotics · Computer Science 2022-04-18 Vlad Sobal , Alfredo Canziani , Nicolas Carion , Kyunghyun Cho , Yann LeCun

Many model-based Visual Odometry (VO) algorithms have been proposed in the past decade, often restricted to the type of camera optics, or the underlying motion manifold observed. We envision robots to be able to learn and perform these…

Robotics · Computer Science 2017-05-30 Sudeep Pillai , John J. Leonard

We propose DOME, a diffusion-based world model that predicts future occupancy frames based on past occupancy observations. The ability of this world model to capture the evolution of the environment is crucial for planning in autonomous…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Songen Gu , Wei Yin , Bu Jin , Xiaoyang Guo , Junming Wang , Haodong Li , Qian Zhang , Xiaoxiao Long

Grounding textual expressions on scene objects from first-person views is a truly demanding capability in developing agents that are aware of their surroundings and behave following intuitive text instructions. Such capability is of…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Shuhei Kurita , Naoki Katsura , Eri Onami

Touch contact and pressure are essential for understanding how humans interact with and manipulate objects, insights which can significantly benefit applications in mixed reality and robotics. However, estimating these interactions from an…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Yiming Zhao , Taein Kwon , Paul Streli , Marc Pollefeys , Christian Holz

Event-based vision sensors offer high time resolution, high dynamic range, and low power consumption, yet event-based vision models lag behind conventional frame-based vision methods. We argue that this gap is partly due to the lack of…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Jens Egholm Pedersen , Dimitris Korakovounis , Jörg Conradt

Unified multimodal generative models aim to integrate image understanding and generation abilities, offering significant advantages in harnessing multimodal corpora, particularly interleaved text-image data. However, existing unified models…

Computer Vision and Pattern Recognition · Computer Science 2025-07-16 Hong Zhang , Zhongjie Duan , Xingjun Wang , Yuze Zhao , Weiyi Lu , Zhipeng Di , Yixuan Xu , Yingda Chen , Yu Zhang

Object tracking is an important functionality of edge video analytic systems and services. Multi-object tracking (MOT) detects the moving objects and tracks their locations frame by frame as real scenes are being captured into a video.…

Computer Vision and Pattern Recognition · Computer Science 2023-09-07 Sanjana Vijay Ganesh , Yanzhao Wu , Gaowen Liu , Ramana Kompella , Ling Liu

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground subject, background…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Jingyun Liang , Jingkai Zhou , Shikai Li , Chenjie Cao , Lei Sun , Yichen Qian , Weihua Chen , Fan Wang

Humans naturally process real-world multimodal information in a full-duplex manner. In artificial intelligence, replicating this capability is essential for advancing model development and deployment, particularly in embodied contexts. The…

Artificial Intelligence · Computer Science 2025-06-03 Yiqun Yao , Xiang Li , Xin Jiang , Xuezhi Fang , Naitong Yu , Aixin Sun , Yequan Wang

Large Multimodal Models (LMMs) have achieved remarkable progress in general-purpose vision--language understanding, yet they remain limited in tasks requiring precise object-level grounding, fine-grained spatial reasoning, and controllable…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Yuqian Yuan , Wenqiao Zhang , Juekai Lin , Yu Zhong , Mingjian Gao , Binhe Yu , Yunqi Cao , Wentong Li , Yueting Zhuang , Beng Chin Ooi

Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire day or more of continuous visual experience. In ultra-long video settings, relevant information is…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Ziyang Wang , Yue Zhang , Shoubin Yu , Ce Zhang , Zengqi Zhao , Jaehong Yoon , Hyunji Lee , Gedas Bertasius , Mohit Bansal

Earth observation (EO) data spans a wide range of spatial, spectral, and temporal resolutions, from high-resolution optical imagery to low resolution multispectral products or radar time series. While recent foundation models have improved…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Nicolas Houdré , Diego Marcos , Hugo Riffaud de Turckheim , Dino Ienco , Laurent Wendling , Camille Kurtz , Sylvain Lobry

This paper addresses the problem of end-to-end self-supervised forecasting of depth and ego motion. Given a sequence of raw images, the aim is to forecast both the geometry and ego-motion using a self supervised photometric loss. The…

Computer Vision and Pattern Recognition · Computer Science 2022-06-16 Houssem Boulahbal , Adrian Voicila , Andrew Comport

We present a scalable self-supervised approach for segmenting feasible vehicle trajectories from monocular images for autonomous driving in complex urban environments. Leveraging large-scale dashcam videos, we treat recorded ego-vehicle…

Robotics · Computer Science 2026-03-25 Tomasz Frelek , Rohan Patil , Akshar Tumu , Henrik I. Christensen

Modeling human behaviors in contextual environments has a wide range of applications in character animation, embodied AI, VR/AR, and robotics. In real-world scenarios, humans frequently interact with the environment and manipulate various…

Computer Vision and Pattern Recognition · Computer Science 2023-09-29 Jiaman Li , Jiajun Wu , C. Karen Liu
‹ Prev 1 8 9 10 Next ›