English
Related papers

Related papers: PhyEdit: Towards Real-World Object Manipulation vi…

200 papers

Corner cases are crucial for training and validating autonomous driving systems, yet collecting them from the real world is often costly and hazardous. Editing objects within captured sensor data offers an effective alternative for…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Jiusi Li , Jackson Jiang , Jinyu Miao , Miao Long , Tuopu Wen , Peijin Jia , Shengxiang Liu , Chunlei Yu , Maolin Liu , Yuzhan Cai , Kun Jiang , Mengmeng Yang , Diange Yang

Diffusion models have significantly improved text-to-image generation, producing high-quality, realistic images from textual descriptions. Beyond generation, object-level image editing remains a challenging problem, requiring precise…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Marco Schouten , Mehmet Onurcan Kaya , Serge Belongie , Dim P. Papadopoulos

Recent advances in image-based 3D human shape estimation have been driven by the significant improvement in representation power afforded by deep neural networks. Although current approaches have demonstrated the potential in real world…

Computer Vision and Pattern Recognition · Computer Science 2020-04-02 Shunsuke Saito , Tomas Simon , Jason Saragih , Hanbyul Joo

Existing single-image 3D indoor scene generators often produce results that look visually plausible but fail to obey real-world physics, limiting their reliability in robotics, embodied AI, and design. To examine this gap, we introduce a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Dongli Wu , Jingyu Hu , Ka-Hei Hui , Xiaobao Wei , Chengwen Luo , Jianqiang Li , Zhengzhe Liu

We present GSEdit, a pipeline for text-guided 3D object editing based on Gaussian Splatting models. Our method enables the editing of the style and appearance of 3D objects without altering their main details, all in a matter of minutes on…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Francesco Palandra , Andrea Sanchietti , Daniele Baieri , Emanuele Rodolà

Realistic image manipulation is challenging because it requires modifying the image appearance in a user-controlled way, while preserving the realism of the result. Unless the user has considerable artistic skill, it is easy to "fall off"…

Computer Vision and Pattern Recognition · Computer Science 2018-12-18 Jun-Yan Zhu , Philipp Krähenbühl , Eli Shechtman , Alexei A. Efros

State-of-the-art object pose estimation methods are prone to generating geometrically infeasible pose hypotheses. This problem is prevalent in dexterous manipulation, where estimated poses often intersect with the robotic hand or are not…

Robotics · Computer Science 2026-03-24 Anil Zeybek , Rhys Newbury , Snehal Dikhale , Nawid Jamali , Soshi Iba , Akansel Cosgun

We introduce MotionEdit, a novel dataset for motion-centric image editing-the task of modifying subject actions and interactions while preserving identity, structure, and physical plausibility. Unlike existing image editing datasets that…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Yixin Wan , Lei Ke , Wenhao Yu , Kai-Wei Chang , Dong Yu

The ultimate goal of many image-based modeling systems is to render photo-realistic novel views of a scene without visible artifacts. Existing evaluation metrics and benchmarks focus mainly on the geometric accuracy of the reconstructed…

Computer Vision and Pattern Recognition · Computer Science 2016-01-27 Michael Waechter , Mate Beljan , Simon Fuhrmann , Nils Moehrle , Johannes Kopf , Michael Goesele

Realistic object interactions are crucial for creating immersive virtual experiences, yet synthesizing realistic 3D object dynamics in response to novel interactions remains a significant challenge. Unlike unconditional or text-conditioned…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Tianyuan Zhang , Hong-Xing Yu , Rundi Wu , Brandon Y. Feng , Changxi Zheng , Noah Snavely , Jiajun Wu , William T. Freeman

Direct mesh editing and deformation are key components in the geometric modeling and animation pipeline. Mesh editing methods are typically framed as optimization problems combining user-specified vertex constraints with a regularizer that…

Graphics · Computer Science 2024-08-05 Tianhao Xie , Eugene Belilovsky , Sudhir Mudur , Tiberiu Popa

Recently, how to achieve precise image editing has attracted increasing attention, especially given the remarkable success of text-to-image generation models. To unify various spatial-aware image editing abilities into one framework, we…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Yueru Jia , Yuhui Yuan , Aosong Cheng , Chuke Wang , Ji Li , Huizhu Jia , Shanghang Zhang

Current geometry-based monocular 3D object detection models can efficiently detect objects by leveraging perspective geometry, but their performance is limited due to the absence of accurate depth information. Though this issue can be…

Computer Vision and Pattern Recognition · Computer Science 2021-07-29 Chenhang He , Jianqiang Huang , Xian-Sheng Hua , Lei Zhang

Machine learning has enabled the development of powerful systems capable of editing images from natural language instructions. However, in many common scenarios it is difficult for users to specify precise image transformations with text…

Artificial Intelligence · Computer Science 2024-02-14 Alec Helbling , Seongmin Lee , Polo Chau

With the rapid advancement of commercial multi-modal models, image editing has garnered significant attention due to its widespread applicability in daily life. Despite impressive progress, existing image editing systems, particularly…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yiran Zhao , Yaoqi Ye , Xiang Liu , Michael Qizhe Shieh , Trung Bui

Recent advances in deep generative modeling have unlocked unprecedented opportunities for video synthesis. In real-world applications, however, users often seek tools to faithfully realize their creative editing intentions with precise and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Yuhao Liu , Tengfei Wang , Fang Liu , Zhenwei Wang , Rynson W. H. Lau

Modern robotic manipulation primarily relies on visual observations in a 2D color space for skill learning but suffers from poor generalization. In contrast, humans, living in a 3D world, depend more on physical properties-such as distance,…

Instruction-based image editing through natural language has emerged as a powerful paradigm for intuitive visual manipulation. While recent models achieve impressive results on single edits, they suffer from severe quality degradation under…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Yucheng Liao , Jiajun Liang , Kaiqian Cui , Baoquan Zhao , Haoran Xie , Wei Liu , Qing Li , Xudong Mao

Being able to edit panoramic images is crucial for creating realistic 360{\deg} visual experiences. However, existing perspective-based image editing methods fail to model the spatial structure of panoramas. Conventional cube-map…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Dong Liang , Yuhao Liu , Jinyuan Jia , Youjun Zhao , Rynson W. H. Lau

Recent advances in image editing, driven by image diffusion models, have shown remarkable progress. However, significant challenges remain, as these models often struggle to follow complex edit instructions accurately and frequently…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Noam Rotstein , Gal Yona , Daniel Silver , Roy Velich , David Bensaïd , Ron Kimmel