English
Related papers

Related papers: Scene-LLM: Extending Language Model for 3D Visual …

200 papers

Recent works have shown how the reasoning capabilities of Large Language Models (LLMs) can be applied to domains beyond natural language processing, such as planning and interaction for robots. These embodied problems require an agent to…

Large Vision and Language Models (LVLMs) have shown strong performance across various vision-language tasks in natural image domains. However, their application to remote sensing (RS) remains underexplored due to significant domain…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Sungjune Park , Yeongyun Kim , Se Yeon Kim , Yong Man Ro

Large language models (LLMs) have unlocked new capabilities of task planning from human instructions. However, prior attempts to apply LLMs to real-world robotic tasks are limited by the lack of grounding in the surrounding scene. In this…

In dynamic open-world environments, autonomous agents often encounter novelties that hinder their ability to find plans to achieve their goals. Specifically, traditional symbolic planners fail to generate plans when the robot's planning…

Robotics · Computer Science 2026-03-13 Hong Lu , Pierrick Lorang , Timothy R. Duggan , Jivko Sinapov , Matthias Scheutz

Recent advances in Large Multimodal Models (LMM) have made it possible for various applications in human-machine interactions. However, developing LMMs that can comprehend, reason, and plan in complex and diverse 3D environments remains a…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Sijin Chen , Xin Chen , Chi Zhang , Mingsheng Li , Gang Yu , Hao Fei , Hongyuan Zhu , Jiayuan Fan , Tao Chen

Recent advances in Large Language Models (LLMs) have helped facilitate exciting progress for robotic planning in real, open-world environments. 3D scene graphs (3DSGs) offer a promising environment representation for grounding such…

Robotics · Computer Science 2024-11-01 Meghan Booker , Grayson Byrd , Bethany Kemp , Aurora Schmidt , Corban Rivera

Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Chengjun Yu , Xuhan Zhu , Chaoqun Du , Pengfei Yu , Wei Zhai , Yang Cao , Zheng-Jun Zha

Understanding 3D medical image volumes is critical in the medical field, yet existing 3D medical convolution and transformer-based self-supervised learning (SSL) methods often lack deep semantic comprehension. Recent advancements in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Qiuhui Chen , Xuancheng Yao , Huping Ye , Yi Hong

We address the new problem of language-guided semantic style transfer of 3D indoor scenes. The input is a 3D indoor scene mesh and several phrases that describe the target scene. Firstly, 3D vertex coordinates are mapped to RGB residues by…

Computer Vision and Pattern Recognition · Computer Science 2022-08-17 Bu Jin , Beiwen Tian , Hao Zhao , Guyue Zhou

Text-to-3D scene generation from natural language is highly desirable for digital content creation. However, existing methods are largely domain-restricted or reliant on predefined spatial relationships, limiting their capacity for…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Jun Luo , Jiaxiang Tang , Ruijie Lu , Gang Zeng

Recent advancements in Generative AI, particularly in Large Language Models (LLMs) and Large Vision-Language Models (LVLMs), offer new possibilities for integrating cognitive planning into robotic systems. In this work, we present a novel…

Robotics · Computer Science 2024-11-06 Arjun P S , Andrew Melnik , Gora Chand Nandi

Large Language Models (LLMs) have demonstrated potential in Vision-and-Language Navigation (VLN) tasks, yet current applications face challenges. While LLMs excel in general conversation scenarios, they struggle with specialized navigation…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Yunzhe Xu , Yiyuan Pan , Zhe Liu , Hesheng Wang

In this paper, we introduce a novel system designed to enhance customer service in the financial and retail sectors through a context-aware 3D virtual agent, utilizing Mixed Reality (MR) and Vision Language Models (VLMs). Our approach…

Human-Computer Interaction · Computer Science 2024-10-17 Cindy Xu , Mengyu Chen , Pranav Deshpande , Elvir Azanli , Runqing Yang , Joseph Ligman

Remarkable progress in 2D Vision-Language Models (VLMs) has spurred interest in extending them to 3D settings for tasks like 3D Question Answering, Dense Captioning, and Visual Grounding. Unlike 2D VLMs that typically process images through…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Haoyuan Li , Yanpeng Zhou , Yufei Gao , Tao Tang , Jianhua Han , Yujie Yuan , Dave Zhenyu Chen , Jiawang Bian , Hang Xu , Xiaodan Liang

Recent advances in metric, semantic, and topological mapping have equipped autonomous robots with semantic concept grounding capabilities to interpret natural language tasks. This work aims to leverage these new capabilities with an…

Vision-Language-Action (VLA) models have emerged as a promising framework for enabling generalist robots capable of perceiving, reasoning, and acting in the real world. These models usually build upon pretrained Vision-Language Models…

Robotics · Computer Science 2025-11-25 Tao Lin , Gen Li , Yilei Zhong , Yanwen Zou , Yuxin Du , Jiting Liu , Encheng Gu , Bo Zhao

Vision-language models (VLMs) show great promise for 3D scene understanding but are mainly applied to indoor spaces or autonomous driving, focusing on low-level tasks like segmentation. This work expands their use to urban-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Valentin Bieri , Marco Zamboni , Nicolas S. Blumer , Qingxuan Chen , Francis Engelmann

Referential grounding in outdoor driving scenes is challenging due to large scene variability, many visually similar objects, and dynamic elements that complicate resolving natural-language references (e.g., "the black car on the right").…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Pranav Saxena , Avigyan Bhattacharya , Ji Zhang , Wenshan Wang

Large Language Models (LLMs) demonstrate substantial potential across a diverse array of domains via request serving. However, as trends continue to push for expanding context sizes, the autoregressive nature of LLMs results in highly…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-07-08 Bin Lin , Chen Zhang , Tao Peng , Hanyu Zhao , Wencong Xiao , Minmin Sun , Anmin Liu , Zhipeng Zhang , Lanbo Li , Xiafei Qiu , Shen Li , Zhigang Ji , Tao Xie , Yong Li , Wei Lin

3D understanding is a key capability for real-world AI assistance. High-quality data plays an important role in driving the development of the 3D understanding community. Current 3D scene understanding datasets often provide geometric and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Zirui Wang , Tao Zhang
‹ Prev 1 8 9 10 Next ›