English
Related papers

Related papers: ScanEdit: Hierarchically-Guided Functional 3D Scan…

200 papers

The rapid advancement of Multimodal Large Language Models (MLLMs) has significantly impacted various multimodal tasks. However, these models face challenges in tasks that require spatial understanding within 3D environments. Efforts to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Duo Zheng , Shijia Huang , Liwei Wang

The ability to provide fine-grained control for generating and editing visual imagery has profound implications for computer vision and its applications. Previous works have explored extending controllability in two directions: instruction…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Shufan Li , Harkanwar Singh , Aditya Grover

Recent advances in visual generative models have enabled high-fidelity image editing guided by human instructions. However, these models often struggle with complex instructions involving combinatorial editing operations or inter-step…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Zilai Zeng , Mingdeng Cao , Zijie Li , Xiaochen Lian , Yichun Shi , Peihao Zhu , Chen Sun , Peng Wang

3D vision-language (VL) reasoning has gained significant attention due to its potential to bridge the 3D physical world with natural language descriptions. Existing approaches typically follow task-specific, highly specialized paradigms.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Hao Liu , Yanni Ma , Yan Liu , Haihong Xiao , Ying He

Generating immersive 3D scenes from texts is a core task in computer vision, crucial for applications in virtual reality and game development. Despite the promise of leveraging 2D diffusion priors, existing methods suffer from spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Jisheng Chu , Wenrui Li , Rui Zhao , Wangmeng Zuo , Shifeng Chen , Xiaopeng Fan

Text-based 3D human motion editing is a critical yet challenging task in computer vision and graphics. While training-free approaches have been explored, the recent release of the MotionFix dataset, which includes source-text-motion…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Zhengyuan Li , Kai Cheng , Anindita Ghosh , Uttaran Bhattacharya , Liangyan Gui , Aniket Bera

We present HumanEdit, a high-quality, human-rewarded dataset specifically designed for instruction-guided image editing, enabling precise and diverse image manipulations through open-form language instructions. Previous large-scale editing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Jinbin Bai , Wei Chow , Ling Yang , Xiangtai Li , Juncheng Li , Hanwang Zhang , Shuicheng Yan

Aiming to link natural language descriptions to specific regions in a 3D scene represented as 3D point clouds, 3D visual grounding is a very fundamental task for human-robot interaction. The recognition errors can significantly impact the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Ziyang Lu , Yunqiang Pei , Guoqing Wang , Yang Yang , Zheng Wang , Heng Tao Shen

We introduce neuralCAD-Edit, the first benchmark for editing 3D CAD models collected from expert CAD engineers. Instead of text conditioning as in prior works, we collect realistic CAD editing requests by capturing videos of professional…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Toby Perrett , Matthew Bouchard , William McCarthy

Applying Gaussian Splatting to perception tasks for 3D scene understanding is becoming increasingly popular. Most existing works primarily focus on rendering 2D feature maps from novel viewpoints, which leads to an imprecise 3D language…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Hao Li , Minghan Qin , Zhengyu Zou , Diqi He , Xinhao Ji , Bohan Li , Bingquan Dai , Dingewn Zhang , Junwei Han

Large Language Models (LLMs) and Vision Language Models (VLMs) have shown impressive reasoning abilities, yet they struggle with spatial understanding and layout consistency when performing fine-grained visual editing. We introduce a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Haoyu Zhen , Xiaolong Li , Yilin Zhao , Han Zhang , Sifei Liu , Kaichun Mo , Chuang Gan , Subhashree Radhakrishnan

Instruction-based image editing aims to modify specific image elements with natural language instructions. However, current models in this domain often struggle to accurately execute complex user instructions, as they are trained on…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Qifan Yu , Wei Chow , Zhongqi Yue , Kaihang Pan , Yang Wu , Xiaoyang Wan , Juncheng Li , Siliang Tang , Hanwang Zhang , Yueting Zhuang

Predicting human gaze scanpaths is crucial for understanding visual attention, with applications in human-computer interaction, autonomous systems, and cognitive robotics. While deep learning models have advanced scanpath prediction, most…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Giuseppe Cartella , Vittorio Cuculo , Alessandro D'Amelio , Marcella Cornia , Giuseppe Boccignone , Rita Cucchiara

Image-text matching plays a central role in bridging the semantic gap between vision and language. The key point to achieve precise visual-semantic alignment lies in capturing the fine-grained cross-modal correspondence between image and…

Computer Vision and Pattern Recognition · Computer Science 2021-06-14 Zhong Ji , Kexin Chen , Haoran Wang

Recent advancements in multi-modal large language models (MLLMs) have shown strong potential for 3D scene understanding. However, existing methods struggle with fine-grained object grounding and contextual reasoning, limiting their ability…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Haifeng Huang , Yilun Chen , Zehan Wang , Jiangmiao Pang , Zhou Zhao

The reconstruction of immersive and realistic 3D scenes holds significant practical importance in various fields of computer vision and computer graphics. Typically, immersive and realistic scenes should be free from obstructions by dynamic…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Zilong Huang , Jun He , Junyan Ye , Lihan Jiang , Weijia Li , Yiping Chen , Ting Han

Three-Dimensional (3D) dense captioning is an emerging vision-language bridging task that aims to generate multiple detailed and accurate descriptions for 3D scenes. It presents significant potential and challenges due to its closer…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Ting Yu , Xiaojun Lin , Shuhui Wang , Weiguo Sheng , Qingming Huang , Jun Yu

Instruction-based image editing holds immense potential for a variety of applications, as it enables users to perform any editing operation using a natural language instruction. However, current models in this domain often struggle with…

Computer Vision and Pattern Recognition · Computer Science 2023-11-17 Shelly Sheynin , Adam Polyak , Uriel Singer , Yuval Kirstain , Amit Zohar , Oron Ashual , Devi Parikh , Yaniv Taigman

Urban modeling is essential for city planning, scene synthesis, and gaming. Existing image-based methods generate diverse layouts but often lack geometric continuity and scalability, while graph-based methods capture structural relations…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Mengyuan Niu , Xinxin Zhuo , Ruizhe Wang , Yuyue Huang , Junyan Yang , Qiao Wang

In addition to color and textural information, geometry provides important cues for 3D scene reconstruction. However, current reconstruction methods only include geometry at the feature level thus not fully exploiting the geometric…

Computer Vision and Pattern Recognition · Computer Science 2024-08-29 Ruihong Yin , Sezer Karaoglu , Theo Gevers