中文
相关论文

相关论文: ArtLLM: Generating Articulated Assets via 3D LLM

200 篇论文

Understanding the 3D motion of articulated objects is essential in robotic scene understanding, mobile manipulation, and motion planning. Prior methods for articulation estimation have primarily focused on controlled settings, assuming…

机器人学 · 计算机科学 2025-09-03 Abdelrhman Werby , Martin Büchner , Adrian Röfer , Chenguang Huang , Wolfram Burgard , Abhinav Valada

We introduce AudioLM, a framework for high-quality audio generation with long-term consistency. AudioLM maps the input audio to a sequence of discrete tokens and casts audio generation as a language modeling task in this representation…

Bridging the gap between complex human instructions and precise 3D object grounding remains a significant challenge in vision and robotics. Existing 3D segmentation methods often struggle to interpret ambiguous, reasoning-based…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Changyue Shi , Minghao Chen , Yiping Mao , Chuxiao Yang , Xinyuan Hu , Jiajun Ding , Zhou Yu

We address the problem of building digital twins of unknown articulated objects from two RGBD scans of the object at different articulation states. We decompose the problem into two stages, each addressing distinct aspects. Our method first…

计算机视觉与模式识别 · 计算机科学 2024-06-10 Yijia Weng , Bowen Wen , Jonathan Tremblay , Valts Blukis , Dieter Fox , Leonidas Guibas , Stan Birchfield

Automatic 3D content creation has gained increasing attention recently, due to its potential in various applications such as video games, film industry, and AR/VR. Recent advancements in diffusion models and multimodal models have notably…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Yitong Wang , Xudong Xu , Li Ma , Haoran Wang , Bo Dai

Graph model generation from natural language description is an important task with many applications in software engineering. With the rise of large language models (LLMs), there is a growing interest in using LLMs for graph model…

软件工程 · 计算机科学 2025-08-04 Boqi Chen , Ou Wei , Bingzhou Zheng , Gunter Mussbacher

We introduce a novel method for real-time animation control and generation on rigged models using natural language input. First, we embed a large language model (LLM) in Unity to output structured texts that can be parsed into diverse and…

Generative AI plays an increasing role during software engineering activities to make them, e.g., more efficient or provide better quality. However, it is often unclear how much benefit LLMs really provide. We concentrate on software…

软件工程 · 计算机科学 2026-01-28 Frank Elberzhager , Matthias Gerbershagen , Joshua Ginkel

In order for robots to operate effectively in homes and workplaces, they must be able to manipulate the articulated objects common within environments built for and by humans. Previous work learns kinematic models that prescribe this…

机器人学 · 计算机科学 2016-07-04 Zhengyang Wu , Mohit Bansal , Matthew R. Walter

Foundational models with billions of parameters which have been trained on large corpora of data have demonstrated non-trivial skills in a variety of domains. However, due to their monolithic structure, it is challenging and expensive to…

We explore a novel method to perceive and manipulate 3D articulated objects that generalizes to enable a robot to articulate unseen classes of objects. We propose a vision-based system that learns to predict the potential motions of the…

机器人学 · 计算机科学 2024-05-03 Ben Eisner , Harry Zhang , David Held

Understanding articulated objects from monocular video is a crucial yet challenging task in robotics and digital twin creation. Existing methods often rely on complex multi-view setups, high-fidelity object scans, or fragile long-term point…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Arslan Artykov , Tom Ravaud , Corentin Sautier , Vincent Lepetit

Robotic vision applications often necessitate a wide range of visual perception tasks, such as object detection, segmentation, and identification. While there have been substantial advances in these individual tasks, integrating specialized…

机器人学 · 计算机科学 2024-02-26 Zijun Long , George Killick , Richard McCreadie , Gerardo Aragon Camarasa

We present BimArt, a novel generative approach for synthesizing 3D bimanual hand interactions with articulated objects. Unlike prior works, we do not rely on a reference grasp, a coarse hand trajectory, or separate modes for grasping and…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Wanyue Zhang , Rishabh Dabral , Vladislav Golyanik , Vasileios Choutas , Eduardo Alvarado , Thabo Beeler , Marc Habermann , Christian Theobalt

Animating an object in 3D often requires an articulated structure, e.g. a kinematic chain or skeleton of the manipulated object with proper skinning weights, to obtain smooth movements and surface deformations. However, existing models that…

计算机视觉与模式识别 · 计算机科学 2023-04-17 Tianshu Kuai , Akash Karthikeyan , Yash Kant , Ashkan Mirzaei , Igor Gilitschenski

Recent advances in 3D AIGC have shown promise in directly creating 3D objects from text and images, offering significant cost savings in animation and product design. However, detailed edit and customization of 3D assets remains a…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Zhangyang Qi , Yunhan Yang , Mengchen Zhang , Long Xing , Xiaoyang Wu , Tong Wu , Dahua Lin , Xihui Liu , Jiaqi Wang , Hengshuang Zhao

As 3D Gaussian Splatting (3DGS) emerges as a leading approach for novel view synthesis and scene reconstruction, its potential in digital asset creation has gained significant attention. An increasing number of asset libraries based on GS…

人机交互 · 计算机科学 2026-04-16 Haotian Mao , Hangyu Zhou , Zhuoxiong Xu , Siyue Wei , Yule Quan , Yan Zhang , Zixuan Guo , Nianchen Deng , Xubo Yang

Active learning (AL) accelerates scientific discovery by prioritizing the most informative experiments, but traditional machine learning (ML) models used in AL suffer from cold-start limitations and domain-specific feature engineering,…

Multi-modal Large Language Models (MLLMs) exhibit impressive capabilities in 2D tasks, yet encounter challenges in discerning the spatial positions, interrelations, and causal logic in scenes when transitioning from 2D to 3D…

计算机视觉与模式识别 · 计算机科学 2025-01-15 Haomiao Xiong , Yunzhi Zhuge , Jiawen Zhu , Lu Zhang , Huchuan Lu

Mixed reality (MR) environments offer embodied spatial interaction, providing intuitive 3D manipulation capabilities that enhance the conceptual design process. Parametric modeling, a powerful and advanced architectural design method,…

人机交互 · 计算机科学 2025-06-09 Ruochen Ji , Lyu Tiangang