English
Related papers

Related papers: Realistic and Controllable 3D Gaussian-Guided Obje…

200 papers

Generative models have achieved significant progress in advancing 2D image editing, demonstrating exceptional precision and realism. However, they often struggle with consistency and object identity preservation due to their inherent…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Yuhuan Xie , Aoxuan Pan , Ming-Xian Lin , Wei Huang , Yi-Hua Huang , Xiaojuan Qi

We present a method for zero-shot, text-driven appearance manipulation in natural images and videos. Given an input image or video and a target text prompt, our goal is to edit the appearance of existing objects (e.g., object's texture) or…

Computer Vision and Pattern Recognition · Computer Science 2022-05-26 Omer Bar-Tal , Dolev Ofri-Amar , Rafail Fridman , Yoni Kasten , Tali Dekel

Understanding the 3D geometry and semantics of driving scenes is critical for safe autonomous driving. Recent advances in 3D occupancy prediction have improved scene representation but often suffer from visual inconsistencies, leading to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Loïck Chambon , Eloi Zablocki , Alexandre Boulch , Mickaël Chen , Matthieu Cord

Generating synthetic images is a useful method for cheaply obtaining labeled data for training computer vision models. However, obtaining accurate 3D models of relevant objects is necessary, and the resulting images often have a gap in…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Bram Vanherle , Brent Zoomers , Jeroen Put , Frank Van Reeth , Nick Michiels

Humans excel at forecasting the future dynamics of a scene given just a single image. Video generation models that can mimic this ability are an essential component for intelligent systems. Recent approaches have improved temporal coherence…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Melonie de Almeida , Daniela Ivanova , Tong Shi , John H. Williamson , Paul Henderson

Advancements in neural implicit representations and differentiable rendering have markedly improved the ability to learn animatable 3D avatars from sparse multi-view RGB videos. However, current methods that map observation space to…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Zichen Tang , Hongyu Yang , Hanchen Zhang , Jiaxin Chen , Di Huang

The realistic reconstruction of street scenes is critical for developing real-world simulators in autonomous driving. Most existing methods rely on object pose annotations, using these poses to reconstruct dynamic objects and move them…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Zipei Ma , Junzhe Jiang , Yurui Chen , Li Zhang

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground subject, background…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Jingyun Liang , Jingkai Zhou , Shikai Li , Chenjie Cao , Lei Sun , Yichen Qian , Weihua Chen , Fan Wang

In this work, we introduce a novel approach for creating controllable dynamics in 3D-generated Gaussians using casually captured reference videos. Our method transfers the motion of objects from reference videos to a variety of generated 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Zhoujie Fu , Jiacheng Wei , Wenhao Shen , Chaoyue Song , Xiaofeng Yang , Fayao Liu , Xulei Yang , Guosheng Lin

The perception of an Autonomous Driving System (ADS) critically depends on relevant, comprehensive, and diverse datasets to ensure its safety while operating in the environment. Field data collection lacks completeness with respect to the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Ali Nouri , Yifei Zhang , Yifan Zhang , Tayssir Bouraffa , Zhennan Fei , Zijian Han , Håkan Sivencrona , Anders Heyden

Object shape provides important information for robotic manipulation; for instance, selecting an effective grasp depends on both the global and local shape of the object of interest, while reaching into clutter requires accurate surface…

Robotics · Computer Science 2019-05-13 Kanrun Huang , Tucker Hermans

The recent Gaussian Splatting achieves high-quality and real-time novel-view synthesis of the 3D scenes. However, it is solely concentrated on the appearance and geometry modeling, while lacking in fine-grained object-level scene…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Mingqiao Ye , Martin Danelljan , Fisher Yu , Lei Ke

Human perception for effective object tracking in 2D video streams arises from the implicit use of prior 3D knowledge and semantic reasoning. In contrast, most generic object tracking (GOT) methods primarily rely on 2D features of the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Shih-Fang Chen , Jun-Cheng Chen , I-Hong Jhuo , Yen-Yu Lin

Recent text-guided generation of individual 3D object has achieved great success using diffusion priors. However, these methods are not suitable for object insertion and replacement tasks as they do not consider the background, leading to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Hanyuan Xiao , Yingshu Chen , Huajian Huang , Haolin Xiong , Jing Yang , Pratusha Prasad , Yajie Zhao

We present I2V3D, a novel framework for animating static images into dynamic videos with precise 3D control, leveraging the strengths of both 3D geometry guidance and advanced generative models. Our approach combines the precision of a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Zhiyuan Zhang , Dongdong Chen , Jing Liao

In the realm of autonomous driving, accurately detecting surrounding obstacles is crucial for effective decision-making. Traditional methods primarily rely on 3D bounding boxes to represent these obstacles, which often fail to capture the…

Robotics · Computer Science 2025-11-18 Chunyong Hu , Qi Luo , Jianyun Xu , Song Wang , Qiang Li , Sheng Yang

Accurate 6D pose estimation of 3D objects is a fundamental task in computer vision, and current research typically predicts the 6D pose by establishing correspondences between 2D image features and 3D model features. However, these methods…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Junbo Li , Weimin Yuan , Yinuo Wang , Yue Zeng , Shihao Shu , Cai Meng , Xiangzhi Bai

We propose a method to enhance 3D Gaussian Splatting (3DGS)~\cite{Kerbl2023}, addressing challenges in initialization, optimization, and density control. Gaussian Splatting is an alternative for rendering realistic images while supporting…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Xingjun Wang , Lianlei Shan

Reconstructing and rendering 3D objects from highly sparse views is of critical importance for promoting applications of 3D vision techniques and improving user experience. However, images from sparse views only contain very limited 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-11-14 Chen Yang , Sikuang Li , Jiemin Fang , Ruofan Liang , Lingxi Xie , Xiaopeng Zhang , Wei Shen , Qi Tian

We present GASPACHO, a method for generating photorealistic, controllable renderings of human-object interactions from multi-view RGB video. Unlike prior work that reconstructs only the human and treats objects as background, GASPACHO…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Aymen Mir , Arthur Moreau , Helisa Dhamo , Zhensong Zhang , Gerard Pons-Moll , Eduardo Pérez-Pellitero