English
Related papers

Related papers: MosaicMem: Hybrid Spatial Memory for Controllable …

200 papers

High-quality 4D reconstruction enables photorealistic and immersive rendering of the dynamic real world. However, unlike static scenes that can be fully captured with a single camera, high-quality dynamic scenes typically require dense…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Weihong Pan , Xiaoyu Zhang , Zhuang Zhang , Zhichao Ye , Nan Wang , Haomin Liu , Guofeng Zhang

Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Aritra Bhowmik , Denis Korzhenkov , Cees G. M. Snoek , Amirhossein Habibian , Mohsen Ghafoorian

Two-stream convolutional networks have shown strong performance in video action recognition tasks. The key idea is to learn spatiotemporal features by fusing convolutional networks spatially and temporally. However, it remains unclear how…

Computer Vision and Pattern Recognition · Computer Science 2019-03-05 Yunbo Wang , Mingsheng Long , Jianmin Wang , Philip S. Yu

Despite recent advancements in text-to-image generation, most existing methods struggle to create images with multiple objects and complex spatial relationships in the 3D world. To tackle this limitation, we introduce a generic AI system,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Yanbo Ding , Shaobin Zhuang , Kunchang Li , Zhengrong Yue , Yu Qiao , Yali Wang

World models have recently gained prominence for action-conditioned visual prediction in complex environments. However, relying on only a few recent observations causes them to lose long-term context. Consequently, within a few steps, the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Nedko Savov , Naser Kazemi , Deheng Zhang , Danda Pani Paudel , Xi Wang , Luc Van Gool

Robotic applications require a comprehensive understanding of the scene. In recent years, neural fields-based approaches that parameterize the entire environment have become popular. These approaches are promising due to their continuous…

Robotics · Computer Science 2024-12-31 Evgenii Kruzhkov , Alena Savinykh , Sven Behnke

Modeling instance-level context and object-object relationships is extremely challenging. It requires reasoning about bounding boxes of different classes, locations \etc. Above all, instance-level spatial reasoning inherently requires…

Computer Vision and Pattern Recognition · Computer Science 2017-04-14 Xinlei Chen , Abhinav Gupta

In this paper, we present a new system for live collaborative dense surface reconstruction. Cooperative robotics, multi participant augmented reality and human-robot interaction are all examples of situations where collaborative mapping can…

Computer Vision and Pattern Recognition · Computer Science 2018-11-22 Louis Gallagher , John B. McDonald

3D human pose reconstruction from single-view camera is a difficult and challenging topic. Many approaches have been proposed, but almost focusing on frame-by-frame independently while inter-frames are highly correlated in a pose sequence.…

Computer Vision and Pattern Recognition · Computer Science 2019-01-11 X. T. Nguyen , T. D. Ngo , T. H. Le

Multi-view generation with camera pose control and prompt-based customization are both essential elements for achieving controllable generative models. However, existing multi-view generation models do not support customization with…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Minjung Shin , Hyunin Cho , Sooyeon Go , Jin-Hwa Kim , Youngjung Uh

Accurate 3D reconstruction from unstructured image collections is a key requirement in applications such as robotics, mapping, and scene understanding. While global Structure from Motion (SfM) techniques rely on full image connectivity and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Muhammad Zeeshan , Umer Zaki , Syed Ahmed Pasha , Zaar Khizar

Multi-frame depth estimation generally achieves high accuracy relying on the multi-view geometric consistency. When applied in dynamic scenes, e.g., autonomous driving, this consistency is usually violated in the dynamic areas, leading to…

Computer Vision and Pattern Recognition · Computer Science 2023-04-19 Rui Li , Dong Gong , Wei Yin , Hao Chen , Yu Zhu , Kaixuan Wang , Xiaozhi Chen , Jinqiu Sun , Yanning Zhang

Monocular 3D human pose estimation poses significant challenges due to the inherent depth ambiguities that arise during the reprojection process from 2D to 3D. Conventional approaches that rely on estimating an over-fit projection matrix…

Computer Vision and Pattern Recognition · Computer Science 2024-01-19 Junkun Jiang , Jie Chen

Humans naturally retain memories of permanent elements, while ephemeral moments often slip through the cracks of memory. This selective retention is crucial for robotic perception, localization, and mapping. To endow robots with this…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Yiming Li , Zehong Wang , Yue Wang , Zhiding Yu , Zan Gojcic , Marco Pavone , Chen Feng , Jose M. Alvarez

Stereo video conversion aims to transform monocular videos into immersive stereo format. Despite the advancements in novel view synthesis, it still remains two major challenges: i) difficulty of achieving high-fidelity and stable results,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Jiale Zhang , Qianxi Jia , Yang Liu , Wei Zhang , Wei Wei , Xin Tian

Deep neural networks have achieved great progress in single-image 3D human reconstruction. However, existing methods still fall short in predicting rare poses. The reason is that most of the current models perform regression based on a…

Computer Vision and Pattern Recognition · Computer Science 2021-01-01 Yu Rong , Ziwei Liu , Chen Change Loy

Constructing compact and informative 3D scene representations is essential for effective embodied exploration and reasoning, especially in complex environments over extended periods. Existing representations, such as object-centric 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Yuncong Yang , Han Yang , Jiachen Zhou , Peihao Chen , Hongxin Zhang , Yilun Du , Chuang Gan

Spatial scene understanding, including monocular depth estimation, is an important problem in various applications, such as robotics and autonomous driving. While improvements in unsupervised monocular depth estimation have potentially…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Hemang Chawla , Arnav Varma , Elahe Arani , Bahram Zonooz

We present MemoryDiorama, a prototype system that introduces augmented memory cues, a concept that extends captured personal media with AI-generated contextual information to enhance autobiographical memory recall. MemoryDiorama transforms…

Human-Computer Interaction · Computer Science 2026-04-09 Keiichi Ihara , Tianle Li , Yasuhisa Shiino , Ryo Suzuki

Metaverse technologies demand accurate, real-time, and immersive modeling on consumer-grade hardware for both non-human perception (e.g., drone/robot/autonomous car navigation) and immersive technologies like AR/VR, requiring both…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Xuqian Ren , Wenjia Wang , Dingding Cai , Tuuli Tuominen , Juho Kannala , Esa Rahtu