中文
相关论文

相关论文: MANSION: Multi-floor lANguage-to-3D Scene generatI…

200 篇论文

Integrating robotic systems in architectural and construction processes is of core interest to increase the efficiency of the building industry. Automated planning for such systems enables design analysis tools and facilitates faster design…

机器人学 · 计算机科学 2021-06-07 Valentin N. Hartmann , Ozgur S. Oguz , Danny Driess , Marc Toussaint , Achim Menges

General-purpose robots coexisting with humans in their environment must learn to relate human language to their perceptions and actions to be useful in a range of daily tasks. Moreover, they need to acquire a diverse repertoire of…

机器人学 · 计算机科学 2022-07-14 Oier Mees , Lukas Hermann , Erick Rosete-Beas , Wolfram Burgard

We present Programmable-Room, a framework which interactively generates and edits a 3D room mesh, given natural language instructions. For precise control of a room's each attribute, we decompose the challenging task into simpler steps such…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Jihyun Kim , Junho Park , Kyeongbo Kong , Suk-Ju Kang

Large language models (LLMs) have demonstrated impressive results in developing generalist planning agents for diverse tasks. However, grounding these plans in expansive, multi-floor, and multi-room environments presents a significant…

机器人学 · 计算机科学 2023-09-29 Krishan Rana , Jesse Haviland , Sourav Garg , Jad Abou-Chakra , Ian Reid , Niko Suenderhauf

We present GuidedSceneGen, a text-to-3D generation framework that produces metrically accurate, globally consistent, and semantically interpretable indoor scenes. Unlike prior text-driven methods that often suffer from geometric drift or…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Stefan Ainetter , Thomas Deixelberger , Edoardo A. Dominici , Philipp Drescher , Konstantinos Vardis , Markus Steinberger

Real-world data collection for embodied agents remains costly and unsafe, calling for scalable, realistic, and simulator-ready 3D environments. However, existing scene-generation systems often rely on rule-based or task-specific pipelines,…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Hongchi Xia , Xuan Li , Zhaoshuo Li , Qianli Ma , Jiashu Xu , Ming-Yu Liu , Yin Cui , Tsung-Yi Lin , Wei-Chiu Ma , Shenlong Wang , Shuran Song , Fangyin Wei

This thesis introduces "Embodied Spatial Intelligence" to address the challenge of creating robots that can perceive and act in the real world based on natural language instructions. To bridge the gap between Large Language Models (LLMs)…

机器人学 · 计算机科学 2025-09-03 Jiading Fang

This paper presents a novel generative approach that outputs 3D indoor environments solely from a textual description of the scene. Current methods often treat scene synthesis as a mere layout prediction task, leading to rooms with…

机器学习 · 计算机科学 2025-02-12 Yao Wei , Matteo Toso , Pietro Morerio , Michael Ying Yang , Alessio Del Bue

We present a large language model (LLM) based system to empower quadrupedal robots with problem-solving abilities for long-horizon tasks beyond short-term motions. Long-horizon tasks for quadrupeds are challenging since they require both a…

机器人学 · 计算机科学 2025-03-20 Yutao Ouyang , Jinhan Li , Yunfei Li , Zhongyu Li , Chao Yu , Koushil Sreenath , Yi Wu

Generating human motions from textual descriptions has gained growing research interest due to its wide range of applications. However, only a few works consider human-scene interactions together with text conditions, which is crucial for…

计算机视觉与模式识别 · 计算机科学 2024-05-14 Zhi Cen , Huaijin Pi , Sida Peng , Zehong Shen , Minghui Yang , Shuai Zhu , Hujun Bao , Xiaowei Zhou

Architectural floor plan design demands joint reasoning over geometry, semantics, and spatial hierarchy, which remains a major challenge for current AI systems. Although recent diffusion and language models improve visual fidelity, they…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Sizhong Qin , Ramon Elias Weber , Xinzheng Lu

To enable robots to comprehend high-level human instructions and perform complex tasks, a key challenge lies in achieving comprehensive scene understanding: interpreting and interacting with the 3D environment in a meaningful way. This…

Vision-and-Language Navigation (VLN) requires the agent to follow language instructions to navigate through 3D environments. One main challenge in VLN is the limited availability of photorealistic training environments, which makes it hard…

计算机视觉与模式识别 · 计算机科学 2023-05-31 Jialu Li , Mohit Bansal

SpatialLM is a large language model designed to process 3D point cloud data and generate structured 3D scene understanding outputs. These outputs include architectural elements like walls, doors, windows, and oriented object boxes with…

计算机视觉与模式识别 · 计算机科学 2025-11-06 Yongsen Mao , Junhao Zhong , Chuan Fang , Jia Zheng , Rui Tang , Hao Zhu , Ping Tan , Zihan Zhou

The recent development in multimodal learning has greatly advanced the research in 3D scene understanding in various real-world tasks such as embodied AI. However, most existing studies are facing two common challenges: 1) they are short of…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Xueying Jiang , Lewei Lu , Ling Shao , Shijian Lu

Prompt-driven scene synthesis allows users to generate complete 3D environments from textual descriptions. Current text-to-scene methods often struggle with complex geometries and object transformations, and tend to show weak adherence to…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Frédéric Berdoz , Luca A. Lanzendörfer , Nick Tuninga , Roger Wattenhofer

Recent advances in text-to-3D scene generation have demonstrated significant potential to transform content creation across multiple industries. Although the research community has made impressive progress in addressing the challenges of…

Generating realistic 3D worlds occupied by moving humans has many applications in games, architecture, and synthetic data creation. But generating such scenes is expensive and labor intensive. Recent work generates human poses and motions…

计算机视觉与模式识别 · 计算机科学 2022-12-09 Hongwei Yi , Chun-Hao P. Huang , Shashank Tripathi , Lea Hering , Justus Thies , Michael J. Black

Humans exhibit a remarkable ability to recognize co-visibility-the 3D regions simultaneously visible in multiple images-even when these images are sparsely distributed across a complex scene. This ability is foundational to 3D vision,…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Chao Chen , Nobel Dang , Juexiao Zhang , Wenkai Sun , Pengfei Zheng , Xuhang He , Yimeng Ye , Jiasheng Zhang , Taarun Srinivas , Chen Feng

City-scale 3D generation is of great importance for the development of embodied intelligence and world models. Existing methods, however, face significant challenges regarding quality, fidelity, and scalability in 3D world generation. Thus,…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Shengyuan Wang , Zhiheng Zheng , Yu Shang , Lixuan He , Yangcheng Yu , Fan Hangyu , Jie Feng , Qingmin Liao , Yong Li