English
Related papers

Related papers: One4D: Unified 4D Generation and Reconstruction vi…

200 papers

While the keypoint-based maps created by sparse monocular simultaneous localisation and mapping (SLAM) systems are useful for camera tracking, dense 3D reconstructions may be desired for many robotic tasks. Solutions involving depth cameras…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Tristan Laidlow , Jan Czarnowski , Stefan Leutenegger

We present Envision3D, a novel method for efficiently generating high-quality 3D content from a single image. Recent methods that extract 3D content from multi-view images generated by diffusion models show great potential. However, it is…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Yatian Pang , Tanghui Jia , Yujun Shi , Zhenyu Tang , Junwu Zhang , Xinhua Cheng , Xing Zhou , Francis E. H. Tay , Li Yuan

3D scene generation has garnered growing attention in recent years and has made significant progress. Generating 4D cities is more challenging than 3D scenes due to the presence of structurally complex, visually diverse objects like…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Haozhe Xie , Zhaoxi Chen , Fangzhou Hong , Ziwei Liu

In this paper, we present DM-Calib, a diffusion-based approach for estimating pinhole camera intrinsic parameters from a single input image. Monocular camera calibration is essential for many 3D vision tasks. However, most existing methods…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Junyuan Deng , Wei Yin , Xiaoyang Guo , Qian Zhang , Xiaotao Hu , Weiqiang Ren , Xiao-Xiao Long , Ping Tan

We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to achieve unification along three axes: the model, the tasks,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Chi Zhang , Jiepeng Wang , Youming Wang , Yuanzhi Liang , Xiaoyan Yang , Zuoxin Li , Haibin Huang , Xuelong Li

Although diffusion-based models can generate high-quality and high-resolution video sequences from textual or image inputs, they lack explicit integration of geometric cues when controlling scene lighting and visual appearance across…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Yuanze Lin , Yi-Wen Chen , Yi-Hsuan Tsai , Ronald Clark , Ming-Hsuan Yang

Decentralized federated learning (DFL) based on low-rank adaptation (LoRA) enables mobile devices with multi-task datasets to collaboratively fine-tune a large language model (LLM) by exchanging locally updated parameters with a subset of…

Machine Learning · Computer Science 2026-02-25 Nuocheng Yang , Sihua Wang , Ouwen Huan , Mingzhe Chen , Tony Q. S. Quek , Changchuan Yin

This work presents Controllable Layer Decomposition (CLD), a method for achieving fine-grained and controllable multi-layer separation of raster images. In practical workflows, designers typically generate and edit each RGBA layer…

Graphics · Computer Science 2025-11-26 Zihao Liu , Zunnan Xu , Shi Shu , Jun Zhou , Ruicheng Zhang , Zhenchao Tang , Xiu Li

Generating high-quality 3D content from text, single images, or sparse view images remains a challenging task with broad applications. Existing methods typically employ multi-view diffusion models to synthesize multi-view images, followed…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Junlin Han , Jianyuan Wang , Andrea Vedaldi , Philip Torr , Filippos Kokkinos

The precise reconstruction of 3D objects from a single RGB image in complex scenes presents a critical challenge in virtual reality, autonomous driving, and robotics. Existing neural implicit 3D representation methods face significant…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Luoxi Zhang , Pragyan Shrestha , Yu Zhou , Chun Xie , Itaru Kitahara

We introduce Drag4D, an interactive framework that integrates object motion control within text-driven 3D scene generation. This framework enables users to define 3D trajectories for the 3D objects generated from a single image, seamlessly…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Minjun Kang , Inkyu Shin , Taeyeop Lee , In So Kweon , Kuk-Jin Yoon

Content-aware layout generation is a critical task in graphic design automation, focused on creating visually appealing arrangements of elements that seamlessly blend with a given background image. The variety of real-world applications…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Zeyang Liu , Le Wang , Sanping Zhou , Yuxuan Wu , Xiaolong Sun , Gang Hua , Haoxiang Li

Synthesizing novel views from monocular videos of dynamic scenes remains a challenging problem. Scene-specific methods that optimize 4D representations with explicit motion priors often break down in highly dynamic regions where multi-view…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Thomas Tanay , Mohammed Brahimi , Michal Nazarczuk , Qingwen Zhang , Sibi Catley-Chandar , Arthur Moreau , Zhensong Zhang , Eduardo Pérez-Pellitero

Video generative models are receiving particular attention given their ability to generate realistic and imaginative frames. Besides, these models are also observed to exhibit strong 3D consistency, significantly enhancing their potential…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Yikai Wang , Xinzhou Wang , Zilong Chen , Zhengyi Wang , Fuchun Sun , Jun Zhu

Video editing using diffusion models has achieved remarkable results in generating high-quality edits for videos. However, current methods often rely on large-scale pretraining, limiting flexibility for specific edits. First-frame-guided…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Chenjian Gao , Lihe Ding , Xin Cai , Zhanpeng Huang , Zibin Wang , Tianfan Xue

Image restoration and enhancement are pivotal for numerous computer vision applications, yet unifying these tasks efficiently remains a significant challenge. Inspired by the iterative refinement capabilities of diffusion models, we propose…

Computer Vision and Pattern Recognition · Computer Science 2024-12-20 Minglong Xue , Jinhong He , Shivakumara Palaiahnakote , Mingliang Zhou

Single-view 3D reconstruction is currently approached from two dominant perspectives: reconstruction of scenes with limited diversity using 3D data supervision or reconstruction of diverse singular objects using large image priors. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Andreea Ardelean , Mert Özer , Bernhard Egger

Unified visual grounding pursues a simple and generic technical route to leverage multi-task data with less task-specific design. The most advanced methods typically present boxes and masks as vertex sequences to model referring detection…

Computer Vision and Pattern Recognition · Computer Science 2023-03-15 Zesen Cheng , Kehan Li , Peng Jin , Xiangyang Ji , Li Yuan , Chang Liu , Jie Chen

Generative models have become increasingly powerful tools for robot motion generation, enabling flexible and multimodal trajectory generation across various tasks. Yet, most existing approaches remain limited in handling multiple types of…

Robotics · Computer Science 2026-01-15 Zewen Yang , Xiaobing Dai , Dian Yu , Zhijun Li , Majid Khadiv , Sandra Hirche , Sami Haddadin

We propose the first framework capable of computing a 4D spatio-temporal grid of video frames and 3D Gaussian particles for each time step using a feed-forward architecture. Our architecture has two main components, a 4D video model and a…