English
Related papers

Related papers: SG2VID: Scene Graphs Enable Fine-Grained Control f…

200 papers

Recent works have shown that Large Language Models (LLMs) can facilitate the grounding of instructions for robotic task planning. Despite this progress, most existing works have primarily focused on utilizing raw images to aid LLMs in…

Robotics · Computer Science 2024-03-12 Zhe Ni , Xiaoxin Deng , Cong Tai , Xinyue Zhu , Qinghongbing Xie , Weihang Huang , Xiang Wu , Long Zeng

Training robots in simulation requires diverse 3D scenes that reflect the specific challenges of downstream tasks. However, scenes that satisfy strict task requirements, such as high-clutter environments with plausible spatial arrangement,…

Robotics · Computer Science 2025-08-27 Nicholas Pfaff , Hongkai Dai , Sergey Zakharov , Shun Iwase , Russ Tedrake

Diffusion-based video generation models have made significant strides, producing outputs with improved visual fidelity, temporal coherence, and user control. These advancements hold great promise for improving surgical education by enabling…

Computer Vision and Pattern Recognition · Computer Science 2024-09-26 Joseph Cho , Samuel Schmidgall , Cyril Zakka , Mrudang Mathur , Dhamanpreet Kaur , Rohan Shad , William Hiesinger

Recent progress in robot learning has been driven by large-scale datasets and powerful visuomotor policy architectures, yet policy robustness remains limited by the substantial cost of collecting diverse demonstrations, particularly for…

Robotics · Computer Science 2026-03-24 Yujie Zhao , Hongwei Fan , Di Chen , Shengcong Chen , Liliang Chen , Xiaoqi Li , Guanghui Ren , Hao Dong

Recent advancements in 3D object generation using diffusion models have achieved remarkable success, but generating realistic 3D urban scenes remains challenging. Existing methods relying solely on 3D diffusion models tend to suffer a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Hanlei Guo , Jiahao Shao , Xinya Chen , Xiyang Tan , Sheng Miao , Yujun Shen , Yiyi Liao

Autonomous laparoscopic camera control must maintain a stable and safe surgical view under rapid tool-tissue interactions while remaining interpretable to surgeons. We present a strategy-grounded framework that couples high-level…

Robotics · Computer Science 2026-02-25 Keyu Zhou , Peisen Xu , Yahao Wu , Jiming Chen , Gaofeng Li , Shunlei Li

Controllable 3D scene generation has extensive applications in virtual reality and interior design, where the generated scenes should exhibit high levels of realism and controllability in terms of geometry. Scene graphs provide a suitable…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Zhifei Yang , Keyang Lu , Chao Zhang , Jiaxing Qi , Hanqi Jiang , Ruifei Ma , Shenglin Yin , Yifan Xu , Mingzhe Xing , Zhen Xiao , Jieyi Long , Guangyao Zhai

4D driving simulation is essential for developing realistic autonomous driving simulators. Despite advancements in existing methods for generating driving scenes, significant challenges remain in view transformation and spatial-temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Lening Wang , Wenzhao Zheng , Dalong Du , Yunpeng Zhang , Yilong Ren , Han Jiang , Zhiyong Cui , Haiyang Yu , Jie Zhou , Jiwen Lu , Shanghang Zhang

Many video workflows benefit from a mixture of user controls with varying granularity, from exact 4D object trajectories and camera paths to coarse text prompts, while existing video generative models are typically trained for fixed input…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Haoyi Duan , Yunzhi Zhang , Yilun Du , Jiajun Wu

Controllable video generation aims to synthesize video content that aligns precisely with user-provided conditions, such as text descriptions and initial images. However, a significant challenge persists in this domain: existing models…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Peng Hu , Yu Gu , Liang Luo , Fuji Ren

3D scene generation conditioned on text prompts has significantly progressed due to the development of 2D diffusion generation models. However, the textual description of 3D scenes is inherently inaccurate and lacks fine-grained control…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Minglin Chen , Longguang Wang , Sheng Ao , Ye Zhang , Kai Xu , Yulan Guo

This paper presents a comprehensive workflow for generating and validating a synthetic dataset designed for robotic surgery instrument segmentation. A 3D reconstruction of the Da Vinci robotic arms was refined and animated in Autodesk Maya…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Giorgio Chiesa , Rossella Borra , Vittorio Lauro , Sabrina De Cillis , Daniele Amparore , Cristian Fiori , Riccardo Renzulli , Marco Grangetto

Reconstructing photorealistic and topology-aware human avatars from monocular videos remains a significant challenge in the fields of computer vision and graphics. While existing 3D human avatar modeling approaches can effectively capture…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Yuze Su , Hongsong Wang , Jie Gui , Liang Wang

Video-to-Video synthesis (Vid2Vid) has achieved remarkable results in generating a photo-realistic video from a sequence of semantic maps. However, this pipeline suffers from high computational cost and long inference latency, which largely…

Computer Vision and Pattern Recognition · Computer Science 2022-07-12 Long Zhuo , Guangcong Wang , Shikai Li , Wayne Wu , Ziwei Liu

Synthesizing photo-realistic visual observations from an ego vehicle's driving trajectory is a critical step towards scalable training of self-driving models. Reconstruction-based methods create 3D scenes from driving logs and synthesize…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Jiageng Mao , Boyi Li , Boris Ivanovic , Yuxiao Chen , Yan Wang , Yurong You , Chaowei Xiao , Danfei Xu , Marco Pavone , Yue Wang

We propose a systematic learning-based approach to the generation of massive quantities of synthetic 3D scenes and arbitrary numbers of photorealistic 2D images thereof, with associated ground truth information, for the purposes of…

Computer Vision and Pattern Recognition · Computer Science 2018-06-21 Chenfanfu Jiang , Siyuan Qi , Yixin Zhu , Siyuan Huang , Jenny Lin , Lap-Fai Yu , Demetri Terzopoulos , Song-Chun Zhu

Diffusion models have enabled remarkably high-quality medical image generation, yet it is challenging to enforce anatomical constraints in generated images. To this end, we propose a diffusion model-based method that supports…

Image and Video Processing · Electrical Eng. & Systems 2024-06-21 Nicholas Konz , Yuwen Chen , Haoyu Dong , Maciej A. Mazurowski

Purpose: Accurate identification of hepatocystic anatomy is critical to preventing surgical complications during laparoscopic cholecystectomy. Deep learning models often struggle with occlusions, long-range dependencies, and capturing the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Yihan Li , Nikhil Churamani , Maria Robu , Imanol Luengo , Danail Stoyanov

Traditional rendering pipelines rely on complex assets, accurate materials and lighting, and substantial computational resources to produce realistic imagery, yet they still face challenges in scalability and realism for populated dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Gonzalo Gomez-Nogales , Yicong Hong , Chongjian Ge , Peiye Zhuang , Marc Comino-Trinidad , Dan Casas , Yi Zhou

In this paper, we propose Img2CAD, the first approach to our knowledge that uses 2D image inputs to generate CAD models with editable parameters. Unlike existing AI methods for 3D model generation using text or image inputs often rely on…

Computer Vision and Pattern Recognition · Computer Science 2024-10-07 Tianrun Chen , Chunan Yu , Yuanqi Hu , Jing Li , Tao Xu , Runlong Cao , Lanyun Zhu , Ying Zang , Yong Zhang , Zejian Li , Linyun Sun