English
Related papers

Related papers: CC-FMO: Camera-Conditioned Zero-Shot Single Image …

200 papers

Using ceiling-mounted cameras (CMCs) for indoor visual capturing opens up a wide range of applications. However, registering CMCs to the target scene layout presents a challenging task. While manual registration with specialized tools is…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Fan Yang , Sosuke Yamao , Ikuo Kusajima , Atsunori Moteki , Shoichi Masui , Shan Jiang

Point cloud registration aligns multiple unposed point clouds into a common reference frame and is a core step for 3D reconstruction and robot localization without initial guess. In this work, we cast registration as conditional generation:…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yue Pan , Tao Sun , Liyuan Zhu , Lucas Nunes , Iro Armeni , Jens Behley , Cyrill Stachniss

Scene-consistent video generation aims to create videos that explore 3D scenes based on a camera trajectory. Previous methods rely on video generation models with external memory for consistency, or iterative 3D reconstruction and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 JiaKui Hu , Jialun Liu , Liying Yang , Xinliang Zhang , Kaiwen Li , Shuang Zeng , Yuanwei Li , Haibin Huang , Chi Zhang , Yanye Lu

Reference-based object composition involves integrating foreground reference image with background scene to produce harmonious fused image. This task becomes particularly challenging in cross-domain scenarios, where models must balance…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Raghu Vamsi Chittersu , Yuvraj Singh Rathore , Pranav Adlinge , Kunal Swami

Text-to-image models are showcasing the impressive ability to create high-quality and diverse generative images. Nevertheless, the transition from freehand sketches to complex scene images remains challenging using diffusion models. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Tianyu Zhang , Xiaoxuan Xie , Xusheng Du , Haoran Xie

Egocentric human motion generation and forecasting with scene-context is crucial for enhancing AR/VR experiences, improving human-robot interaction, advancing assistive technologies, and enabling adaptive healthcare solutions by accurately…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Chaitanya Patel , Hiroki Nakamura , Yuta Kyuragi , Kazuki Kozuka , Juan Carlos Niebles , Ehsan Adeli

Humans excel at forecasting the future dynamics of a scene given just a single image. Video generation models that can mimic this ability are an essential component for intelligent systems. Recent approaches have improved temporal coherence…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Melonie de Almeida , Daniela Ivanova , Tong Shi , John H. Williamson , Paul Henderson

We report Zero123++, an image-conditioned diffusion model for generating 3D-consistent multi-view images from a single input view. To take full advantage of pretrained 2D generative priors, we develop various conditioning and training…

Computer Vision and Pattern Recognition · Computer Science 2023-10-24 Ruoxi Shi , Hansheng Chen , Zhuoyang Zhang , Minghua Liu , Chao Xu , Xinyue Wei , Linghao Chen , Chong Zeng , Hao Su

Image caption generation is one of the most challenging problems at the intersection of vision and language domains. In this work, we propose a realistic captioning task where the input scenes may incorporate visual objects with no…

Computer Vision and Pattern Recognition · Computer Science 2022-07-04 Berkan Demirel , Ramazan Gokberk Cinbis

In this work, we consider the problem of estimating the 3D position of multiple humans in a scene as well as their body shape and articulation from a single RGB video recorded with a static camera. In contrast to expensive marker-based or…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Diogo Luvizon , Marc Habermann , Vladislav Golyanik , Adam Kortylewski , Christian Theobalt

Reconstructing physically stable 3D scenes from a single RGB image enables casual images to be converted into simulation-ready digital assets for applications such as immersive interaction and content creation. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Xiaoxuan Ma , Jiashun Wang , Nicolas Ugrinovic , Yehonathan Litman , Kris Kitani

Recent camera-based 3D semantic scene completion (SSC) methods have increasingly explored leveraging temporal cues to enrich the features of the current frame. However, while these approaches primarily focus on enhancing in-frame regions,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Jongseong Bae , Junwoo Ha , Jinnyeong Heo , Yeongin Lee , Ha Young Kim

Existing diffusion-based 3D scene generation methods primarily operate in 2D image/video latent spaces, which makes maintaining cross-view appearance and geometric consistency inherently challenging. To bridge this gap, we present OneWorld,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Sensen Gao , Zhaoqing Wang , Qihang Cao , Dongdong Yu , Changhu Wang , Tongliang Liu , Mingming Gong , Jiawang Bian

We present Fillerbuster, a unified model that completes unknown regions of a 3D scene with a multi-view latent diffusion transformer. Casual captures are often sparse and miss surrounding content behind objects or above the scene. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Ethan Weber , Norman Müller , Yash Kant , Vasu Agrawal , Michael Zollhöfer , Angjoo Kanazawa , Christian Richardt

Robots typically possess sensors of different modalities, such as colour cameras, inertial measurement units, and 3D laser scanners. Often, solving a particular problem becomes easier when more than one modality is used. However, while…

Computer Vision and Pattern Recognition · Computer Science 2017-01-10 Charika De Alvis , Lionel Ott , Fabio Ramos

With the rapid advancement of autonomous driving technology, a lack of data has become a major obstacle to enhancing perception model accuracy. Researchers are now exploring controllable data generation using world models to diversify…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Xinqing Li , Ruiqi Song , Qingyu Xie , Ye Wu , Nanxin Zeng , Yunfeng Ai

Scene modeling is very crucial for robots that need to perceive, reason about and manipulate the objects in their environments. In this paper, we adapt and extend Boltzmann Machines (BMs) for contextualized scene modeling. Although there…

Robotics · Computer Science 2018-12-20 Ilker Bozcan , Sinan Kalkan

Automatic scene generation is an essential area of research with applications in robotics, recreation, visual representation, training and simulation, education, and more. This survey provides a comprehensive review of the current…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Awal Ahmed Fime , Saifuddin Mahmud , Arpita Das , Md. Sunzidul Islam , Hong-Hoon Kim

Recent advances in text-to-3D scene generation have demonstrated significant potential to transform content creation across multiple industries. Although the research community has made impressive progress in addressing the challenges of…

Few-shot image classification remains a critical challenge in the field of computer vision, particularly in data-scarce environments. Existing methods typically rely on pre-trained visual-language models, such as CLIP. However, due to the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Xi Yang , Pai Peng , Wulin Xie , Xiaohuan Lu , Jie Wen