English
Related papers

Related papers: InsertAnywhere: Bridging 4D Scene Geometry and Dif…

200 papers

Gaussian Splatting has become a popular technique for various 3D Computer Vision tasks, including novel view synthesis, scene reconstruction, and dynamic scene rendering. However, the challenge of natural-looking object insertion, where the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Vsevolod Skorokhodov , Nikita Durasov , Pascal Fua

Video editing increasingly demands the ability to incorporate specific real-world instances into existing footage, yet current approaches fundamentally fail to capture the unique visual characteristics of particular subjects and ensure…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Shaobin Zhuang , Zhipeng Huang , Binxin Yang , Ying Zhang , Fangyikang Wang , Canmiao Fu , Chong Sun , Zheng-Jun Zha , Chen Li , Yali Wang

Despite significant advancements in video generation, inserting a given object into videos remains a challenging task. The difficulty lies in preserving the appearance details of the reference object and accurately modeling coherent motions…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Yuanpeng Tu , Hao Luo , Xi Chen , Sihui Ji , Xiang Bai , Hengshuang Zhao

Video inpainting fills in corrupted video content with plausible replacements. While recent advances in endoscopic video inpainting have shown potential for enhancing the quality of endoscopic videos, they mainly repair 2D visual…

Image and Video Processing · Electrical Eng. & Systems 2024-07-04 Francis Xiatian Zhang , Shuang Chen , Xianghua Xie , Hubert P. H. Shum

Open-vocabulary 3D object detection has gained significant interest due to its critical applications in autonomous driving and embodied AI. Existing detection methods, whether offline or online, typically rely on dense point cloud…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Yuqing Lan , Chenyang Zhu , Zhirui Gao , Jiazhao Zhang , Yihan Cao , Renjiao Yi , Yijie Wang , Kai Xu

Creative processes such as painting often involve creating different components of an image one by one. Can we build a computational model to perform this task? Prior works often fail by making global changes to the image, inserting objects…

Computer Vision and Pattern Recognition · Computer Science 2024-12-25 Alper Canberk , Maksym Bondarenko , Ege Ozguroglu , Ruoshi Liu , Carl Vondrick

Text-to-image diffusion models have attracted considerable interest due to their wide applicability across diverse fields. However, challenges persist in creating controllable models for personalized object generation. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2023-06-30 Yuheng Li , Haotian Liu , Yangming Wen , Yong Jae Lee

Removing undesired concepts from large-scale text-to-image (T2I) and text-to-video (T2V) diffusion models while preserving overall generative quality remains a major challenge, particularly as modern models such as Stable Diffusion v3,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zhaoxin Fan , Nanxiang Jiang , Daiheng Gao , Shiji Zhou , Wenjun Wu

Large scale text-guided diffusion models have garnered significant attention due to their ability to synthesize diverse images that convey complex visual concepts. This generative power has more recently been leveraged to perform text-to-3D…

Computer Vision and Pattern Recognition · Computer Science 2023-09-20 Etai Sella , Gal Fiebelman , Peter Hedman , Hadar Averbuch-Elor

The field of generative models has recently witnessed significant progress, with diffusion models showing remarkable performance in image generation. In light of this success, there is a growing interest in exploring the application of…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Ariel Lapid , Idan Achituve , Lior Bracha , Ethan Fetaya

Generating pose-aligned 3D objects is challenging due to the spatial mismatches and transformation ambiguities inherent in decoupled canonical-then-rotate paradigms. To this end, we introduce Pose-Aware Diffusion (PAD), a novel end-to-end…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Zihan Zhou , Luxi Chen , Jingzhi Zhou , Yuhao Wan , Min Zhao , Baoyu Fan , Chongxuan Li

The landscape of AI video generation is undergoing a pivotal shift: moving beyond general generation - which relies on exhaustive prompt-engineering and "cherry-picking" - towards fine-grained, controllable generation and high-fidelity…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Xiangbo Gao , Renjie Li , Xinghao Chen , Yuheng Wu , Suofei Feng , Qing Yin , Zhengzhong Tu

Creating editable videos that depict complex interactions between multiple objects in various artistic styles has long been a challenging task in filmmaking. Progress is often hampered by the scarcity of data sets that contain paired text…

Computer Vision and Pattern Recognition · Computer Science 2024-08-30 Anisha Jain

Visual servoing, the method of controlling robot motion through feedback from visual sensors, has seen significant advancements with the integration of optical flow-based methods. However, its application remains limited by inherent…

Generating high-quality textures for 3D scenes is crucial for applications in interior design, gaming, and augmented/virtual reality (AR/VR). Although recent advancements in 3D generative models have enhanced content creation, significant…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Yunfan Zhang , Zhiwei Xiong , Zhiqi Shen , Guosheng Lin , Hao Wang , Nicolas Vun

Diffusion models have marked a significant milestone in the enhancement of image and video generation technologies. However, generating videos that precisely retain the shape and location of moving objects such as robots remains a…

Robotics · Computer Science 2024-07-04 Peng Wang , Zhihao Guo , Abdul Latheef Sait , Minh Huy Pham

In this work, we address the lack of 3D understanding of generative neural networks by introducing a persistent 3D feature embedding for view synthesis. To this end, we propose DeepVoxels, a learned representation that encodes the…

Computer Vision and Pattern Recognition · Computer Science 2019-04-12 Vincent Sitzmann , Justus Thies , Felix Heide , Matthias Nießner , Gordon Wetzstein , Michael Zollhöfer

A fundamental challenge in text-to-3D face generation is achieving high-quality geometry. The core difficulty lies in the arbitrary and intricate distribution of vertices in 3D space, making it challenging for existing models to establish…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Junyi Zhang , Yiming Wang , Yunhong Lu , Qichao Wang , Wenzhe Qian , Xiaoyin Xu , David Gu , Min Zhang

Video virtual try-on aims to naturally fit a garment to a target person in consecutive video frames. It is a challenging task, on the one hand, the output video should be in good spatial-temporal consistency, on the other hand, the details…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Cheng Zou , Senlin Cheng , Bolei Xu , Dandan Zheng , Xiaobo Li , Jingdong Chen , Ming Yang

Predictive world models that simulate future observations under explicit camera control are fundamental to interactive AI. Despite rapid advances, current systems lack spatial persistence: they fail to maintain stable scene structures over…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Chendong Xiang , Jiajun Liu , Jintao Zhang , Xiao Yang , Zhengwei Fang , Shizun Wang , Zijun Wang , Yingtian Zou , Hang Su , Jun Zhu