English
Related papers

Related papers: SceneSmith: Agentic Generation of Simulation-Ready…

200 papers

This paper presents SceneCut, a novel approach to jointly discover previously unseen objects and non-object surfaces using a single RGB-D image. SceneCut's joint reasoning over scene semantics and geometry allows a robot to detect and…

Computer Vision and Pattern Recognition · Computer Science 2018-05-25 Trung Pham , Thanh-Toan Do , Niko Sünderhauf , Ian Reid

Humans exhibit incredibly high levels of multi-modal understanding - combining visual cues with read, or heard knowledge comes easy to us and allows for very accurate interaction with the surrounding environment. Various simulation…

Robotics · Computer Science 2022-06-22 Michal Nazarczuk , Tony Ng , Krystian Mikolajczyk

Real-world tool-using agents operate over long-horizon workflows with recurring structure and diverse demands, where effective behavior requires not only invoking atomic tools but also abstracting, and reusing higher-level tool…

Text-driven 3D indoor scene generation is useful for gaming, the film industry, and AR/VR applications. However, existing methods cannot faithfully capture the room layout, nor do they allow flexible editing of individual objects in the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Chuan Fang , Yuan Dong , Kunming Luo , Xiaotao Hu , Rakesh Shrestha , Ping Tan

Embodiment is an important characteristic for all intelligent agents (creatures and robots), while existing scene description tasks mainly focus on analyzing images passively and the semantic understanding of the scenario is separated from…

Robotics · Computer Science 2020-05-08 Sinan Tan , Huaping Liu , Di Guo , Xinyu Zhang , Fuchun Sun

Designing 3D scenes is traditionally a challenging task that demands both artistic expertise and proficiency with complex software. Recent advances in text-to-3D generation have greatly simplified this process by letting users create scenes…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Zeqi Gu , Yin Cui , Zhaoshuo Li , Fangyin Wei , Yunhao Ge , Jinwei Gu , Ming-Yu Liu , Abe Davis , Yifan Ding

As more applications of large language models (LLMs) for 3D content for immersive environments emerge, it is crucial to study user behaviour to identify interaction patterns and potential barriers to guide the future design of immersive…

Human-Computer Interaction · Computer Science 2026-04-09 Junlong Chen , Jens Grubert , Per Ola Kristensson

Real-world robotic tasks are long-horizon and often span multiple floors, demanding rich spatial reasoning. However, existing embodied benchmarks are largely confined to single-floor in-house environments, failing to reflect the complexity…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Lirong Che , Shuo Wen , Shan Huang , Chuang Wang , Yuzhe Yang , Gregory Dudek , Xueqian Wang , Jian Su

Seamless Human-Robot Interaction is the ultimate goal of developing service robotic systems. For this, the robotic agents have to understand their surroundings to better complete a given task. Semantic scene understanding allows a robotic…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 Muraleekrishna Gopinathan , Giang Truong , Jumana Abu-Khalaf

Virtual environments are essential to AI agent research. Existing environments for LLM agent research typically focus on either physical task solving or social simulation, with the former oversimplifying agent individuality and social…

Multiagent Systems · Computer Science 2025-06-17 Dekun Wu , Frederik Brudy , Bang Liu , Yi Wang

We consider the problem of generating realistic traffic scenes automatically. Existing methods typically insert actors into the scene according to a set of hand-crafted heuristics and are limited in their ability to model the true…

Computer Vision and Pattern Recognition · Computer Science 2021-01-19 Shuhan Tan , Kelvin Wong , Shenlong Wang , Sivabalan Manivasagam , Mengye Ren , Raquel Urtasun

3D indoor scenes are widely used in computer graphics, with applications ranging from interior design to gaming to virtual and augmented reality. They also contain rich information, including room layout, as well as furniture type,…

Graphics · Computer Science 2023-02-22 Lin Gao , Jia-Mu Sun , Kaichun Mo , Yu-Kun Lai , Leonidas J. Guibas , Jie Yang

Modern vision-language models (VLMs) are expected to have abilities of spatial reasoning with diverse scene complexities, but evaluating such abilities is difficult due to the lack of benchmarks that are not only diverse and scalable but…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Haoming Wang , Qiyao Xue , Wei Gao

The growing complexity of urban mobility systems has made traffic simulation indispensable for evidence-based transportation planning and policy evaluation. However, despite the analytical capabilities of platforms such as the Simulation of…

Human-Computer Interaction · Computer Science 2025-11-11 Minwoo Jeong , Jeeyun Chang , Yoonjin Yoon

Autonomous 3D indoor scene synthesis breaks down in non-convex rooms with tightly coupled spatial constraints. Data-driven generators lack topological priors for long-horizon planning, while iterative agents fragment semantics and become…

Robotics · Computer Science 2026-05-05 Meisheng Zhang , Shizhao Sun , Yang Zhao , Ziyuan Liu , Zhijun Gao , Jiang Bian

Scene simulation in autonomous driving has gained significant attention because of its huge potential for generating customized data. However, existing editable scene simulation approaches face limitations in terms of user interaction…

Computer Vision and Pattern Recognition · Computer Science 2024-06-27 Yuxi Wei , Zi Wang , Yifan Lu , Chenxin Xu , Changxing Liu , Hao Zhao , Siheng Chen , Yanfeng Wang

Advancements in large language models offer strong potential for enhancing virtual simulated patients (VSPs) in medical education by providing scalable alternatives to resource-intensive traditional methods. However, current VSPs often…

Computation and Language · Computer Science 2025-12-23 Victor De Marez , Jens Van Nooten , Luna De Bruyne , Walter Daelemans

We introduce the Scene Language, a visual scene representation that concisely and precisely describes the structure, semantics, and identity of visual scenes. It represents a scene with three key components: a program that specifies the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Yunzhi Zhang , Zizhang Li , Matt Zhou , Shangzhe Wu , Jiajun Wu

With the recent rise of Large Language Models (LLMs), Vision-Language Models (VLMs), and other general foundation models, there is growing potential for multimodal, multi-task embodied agents that can operate in diverse environments given…

Robotics · Computer Science 2024-11-07 Haochen Zhang , Nader Zantout , Pujith Kachana , Zongyuan Wu , Ji Zhang , Wenshan Wang

Sketching serves as a versatile tool for externalizing ideas, enabling rapid exploration and visual communication that spans various disciplines. While artificial systems have driven substantial advances in content creation and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Yael Vinker , Tamar Rott Shaham , Kristine Zheng , Alex Zhao , Judith E Fan , Antonio Torralba