中文
相关论文

相关论文: ArtLLM: Generating Articulated Assets via 3D LLM

200 篇论文

Generating and editing a 3D scene guided by natural language poses a challenge, primarily due to the complexity of specifying the positional relations and volumetric changes within the 3D space. Recent advancements in Large Language Models…

计算机视觉与模式识别 · 计算机科学 2023-05-26 Yiqi Lin , Hao Wu , Ruichen Wang , Haonan Lu , Xiaodong Lin , Hui Xiong , Lin Wang

Simulation is a central tool for scalable robot learning, but its effectiveness depends on the quality of object assets. While modern 3D datasets provide rich geometric and kinematic representations, they typically lack the physical…

机器人学 · 计算机科学 2026-05-20 Anh-Quan Pham

Converting static 3D meshes into interactable articulated assets is crucial for embodied AI and robotic simulation. However, existing zero-shot pipelines struggle with complex assets due to a critical lack of physical grounding.…

机器人学 · 计算机科学 2026-03-16 WenBo Xu , Liu Liu , Li Zhang , Dan Guo , RuoNan Liu

Enabling Large Language Models (LLMs) to comprehend the 3D physical world remains a significant challenge. Due to the lack of large-scale 3D-text pair datasets, the success of LLMs has yet to be replicated in 3D understanding. In this…

计算机视觉与模式识别 · 计算机科学 2025-05-23 Yuan Tang , Xu Han , Xianzhi Li , Qiao Yu , Jinfeng Xu , Yixue Hao , Long Hu , Min Chen

3D scene graphs have empowered robots with semantic understanding for navigation and planning. However, current functional scene graphs primarily focus on static element detection, lacking the actionable kinematic information required for…

机器人学 · 计算机科学 2026-03-24 Qiuyi Gu , Yuze Sheng , Jincheng Yu , Jiahao Tang , Xiaolong Shan , Zhaoyang Shen , Tinghao Yi , Xiaodan Liang , Xinlei Chen , Yu Wang

Animatable 3D human reconstruction from a single image is a challenging problem due to the ambiguity in decoupling geometry, appearance, and deformation. Recent advances in 3D human reconstruction mainly focus on static human modeling, and…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Lingteng Qiu , Xiaodong Gu , Peihao Li , Qi Zuo , Weichao Shen , Junfei Zhang , Kejie Qiu , Weihao Yuan , Guanying Chen , Zilong Dong , Liefeng Bo

Object-centric reconstruction seeks to recover the 3D structure of a scene through composition of independent objects. While this independence can simplify modeling, it discards strong signals that could improve reconstruction, notably…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Qirui Wu , Yawar Siddiqui , Duncan Frost , Samir Aroudj , Armen Avetisyan , Richard Newcombe , Angel X. Chang , Jakob Engel , Henry Howard-Jenkins

Large language models (LLMs) have shown impressive performance on language tasks but face challenges when deployed on resource-constrained devices due to their extensive parameters and reliance on dense multiplications, resulting in high…

The emergence of large language models (LLMs) has revolutionized machine learning and related fields, showcasing remarkable abilities in comprehending, generating, and manipulating human language. However, their conventional usage through…

Crystal structure generation is fundamental to materials science, enabling the discovery of novel materials with desired properties. While existing approaches leverage Large Language Models (LLMs) through extensive fine-tuning on materials…

Location-based services play an critical role in improving the quality of our daily lives. Despite the proliferation of numerous specialized AI models within spatio-temporal context of location-based services, these models struggle to…

机器学习 · 计算机科学 2024-06-19 Yue Jiang , Qin Chao , Yile Chen , Xiucheng Li , Shuai Liu , Gao Cong

This paper presents ShapeLLM, the first 3D Multimodal Large Language Model (LLM) designed for embodied interaction, exploring a universal 3D object understanding with 3D point clouds and languages. ShapeLLM is built upon an improved 3D…

计算机视觉与模式识别 · 计算机科学 2024-07-15 Zekun Qi , Runpei Dong , Shaochen Zhang , Haoran Geng , Chunrui Han , Zheng Ge , Li Yi , Kaisheng Ma

Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced performance on 2D visual tasks. However, improving their spatial intelligence remains a challenge. Existing 3D MLLMs always rely on additional 3D or…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Diankun Wu , Fangfu Liu , Yi-Hsin Hung , Yueqi Duan

Multi-purpose Large Language Models (LLMs), a subset of generative Artificial Intelligence (AI), have recently made significant progress. While expectations for LLMs to assist systems engineering (SE) tasks are paramount; the…

计算与语言 · 计算机科学 2025-02-17 Taylan G. Topcu , Mohammed Husain , Max Ofsa , Paul Wach

Discovering materials with desirable properties in an efficient way remains a significant problem in materials science. Many studies have tackled this problem by using different sets of information available about the materials. Among them,…

材料科学 · 物理学 2025-03-04 Onur Boyar , Indra Priyadarsini , Seiji Takeda , Lisa Hamada

Despite large-scale pretraining endowing models with language and vision reasoning capabilities, improving their spatial reasoning capability remains challenging due to the lack of data grounded in the 3D world. While it is possible for…

Reconstructing articulated objects prevalent in daily environments is crucial for applications in augmented/virtual reality and robotics. However, existing methods face scalability limitations (requiring 3D supervision or costly…

Manipulating articulated objects with robotic arms is challenging due to the complex kinematic structure, which requires precise part segmentation for efficient manipulation. In this work, we introduce a novel superpoint-based perception…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Qiaojun Yu , Ce Hao , Xibin Yuan , Li Zhang , Liu Liu , Yukang Huo , Rohit Agarwal , Cewu Lu

We address the challenge of generating 3D articulated objects in a controllable fashion. Currently, modeling articulated 3D objects is either achieved through laborious manual authoring, or using methods from prior work that are hard to…

计算机视觉与模式识别 · 计算机科学 2024-03-21 Jiayi Liu , Hou In Ivan Tam , Ali Mahdavi-Amiri , Manolis Savva

We present a learning method for predicting animation skeletons for input 3D models of articulated characters. In contrast to previous approaches that fit pre-defined skeleton templates or predict fixed sets of joints, our method produces…

计算机视觉与模式识别 · 计算机科学 2019-08-23 Zhan Xu , Yang Zhou , Evangelos Kalogerakis , Karan Singh