English
Related papers

Related papers: ArtiWorld: LLM-Driven Articulation of 3D Objects i…

200 papers

Enabling Large Language Models (LLMs) to interact with 3D environments is challenging. Existing approaches extract point clouds either from ground truth (GT) geometry or 3D scenes reconstructed by auxiliary models. Text-image aligned 2D…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Tao Chu , Pan Zhang , Xiaoyi Dong , Yuhang Zang , Qiong Liu , Jiaqi Wang

Cities, as the essential environment of human life, encompass diverse physical elements such as buildings, roads and vegetation, which continuously interact with dynamic entities like people and vehicles. Crafting realistic, interactive 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Yu Shang , Yuming Lin , Yu Zheng , Hangyu Fan , Jingtao Ding , Jie Feng , Jiansheng Chen , Li Tian , Yong Li

This paper focuses on the challenging problem of 3D pose estimation of a diverse spectrum of articulated objects from single depth images. A novel structured prediction approach is considered, where 3D poses are represented as skeletal…

Computer Vision and Pattern Recognition · Computer Science 2016-12-05 Yu Zhang , Chi Xu , Li Cheng

The increasing demand for augmented reality and robotics is driving the need for articulated object reconstruction with high scalability. However, existing settings for reconstructing from discrete articulation states or casual monocular…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Hang Dai , Hongwei Fan , Han Zhang , Duojin Wu , Jiyao Zhang , Hao Dong

Realistic simulation of dynamic scenes requires accurately capturing diverse material properties and modeling complex object interactions grounded in physical principles. However, existing methods are constrained to basic material types…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Zhuoman Liu , Weicai Ye , Yan Luximon , Pengfei Wan , Di Zhang

3D generation from natural language offers significant potential to reduce expert manual modeling efforts and enhance accessibility to 3D assets. However, existing methods often yield unstructured meshes and exhibit poor interactivity,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Shuyuan Zhang , Chenhan Jiang , Zuoou Li , Jiankang Deng

3D object grounding localizes referred objects in a 3D scene from natural language. Unified instance-centric 3D-LLMs aim to solve grounding together with dialog, QA, and captioning, yet many rely on a single pointer-style grounding decision…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Jiawei Li , Ziyi Liu , Weijie Shi , Long Chen , Jiajie Xu , Xiaofang Zhou

Perceiving and interacting with 3D articulated objects, such as cabinets, doors, and faucets, pose particular challenges for future home-assistant robots performing daily tasks in human environments. Besides parsing the articulated parts…

Computer Vision and Pattern Recognition · Computer Science 2023-05-05 Yian Wang , Ruihai Wu , Kaichun Mo , Jiaqi Ke , Qingnan Fan , Leonidas Guibas , Hao Dong

Autonomous vehicles operate in highly dynamic environments necessitating an accurate assessment of which aspects of a scene are moving and where they are moving to. A popular approach to 3D motion estimation, termed scene flow, is to employ…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Philipp Jund , Chris Sweeney , Nichola Abdo , Zhifeng Chen , Jonathon Shlens

Understanding 3D scenes in open-world settings poses fundamental challenges for vision and robotics, particularly due to the limitations of closed-vocabulary supervision and static annotations. To address this, we propose a unified…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Fei Yu , Quan Deng , Shengeng Tang , Yuehua Li , Lechao Cheng

Articulated objects are ubiquitous in daily life. In this paper, we present DexSim2Real$^{2}$, a novel framework for goal-conditioned articulated object manipulation. The core of our framework is constructing an explicit world model of…

Robotics · Computer Science 2025-07-15 Taoran Jiang , Yixuan Guan , Liqian Ma , Jing Xu , Jiaojiao Meng , Weihang Chen , Zecui Zeng , Lusong Li , Dan Wu , Rui Chen

Robots operating in human environments must be able to rearrange objects into semantically-meaningful configurations, even if these objects are previously unseen. In this work, we focus on the problem of building physically-valid structures…

Robotics · Computer Science 2023-04-26 Weiyu Liu , Yilun Du , Tucker Hermans , Sonia Chernova , Chris Paxton

Articulated objects are prevalent in daily life. Interactable digital twins of such objects have numerous applications in embodied AI and robotics. Unfortunately, current methods to digitize articulated real-world objects require carefully…

Graphics · Computer Science 2025-11-18 Weikun Peng , Jun Lv , Cewu Lu , Manolis Savva

On-the-fly 3D reconstruction from monocular image sequences is a long-standing challenge in computer vision, critical for applications such as real-to-sim, AR/VR, and robotics. Existing methods face a major tradeoff: per-scene optimization…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Guanghao Li , Kerui Ren , Linning Xu , Zhewen Zheng , Changjian Jiang , Xin Gao , Bo Dai , Jian Pu , Mulin Yu , Jiangmiao Pang

Recent years have produced a variety of learning based methods in the context of computer vision and robotics. Most of the recently proposed methods are based on deep learning, which require very large amounts of data compared to…

Computer Vision and Pattern Recognition · Computer Science 2024-10-28 Kenneth Blomqvist , Julius Hietala

Large Language Models (LLMs) motivate generative agent simulation (e.g., AI Town) to create a ``dynamic world'', holding immense value across entertainment and research. However, for non-experts, especially those without programming skills,…

Human-Computer Interaction · Computer Science 2026-01-30 Jianwen Sun , Yukang Feng , Kaining Ying , Chuanhao Li , Zizhen Li , Fanrui Zhang , Jiaxin Ai , Yifan Chang , Yu Dai , Yifei Huang , Kaipeng Zhang

Articulated object manipulation poses a unique challenge compared to rigid object manipulation as the object itself represents a dynamic environment. In this work, we present a novel RL-based pipeline equipped with variable impedance…

Robotics · Computer Science 2025-02-21 Tan-Dzung Do , Nandiraju Gireesh , Jilong Wang , He Wang

Reasoning segmentation aims to segment target objects in complex scenes based on human intent and spatial reasoning. While recent multimodal large language models (MLLMs) have demonstrated impressive 2D image reasoning segmentation,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Jiaxin Huang , Runnan Chen , Ziwen Li , Zhengqing Gao , Xiao He , Yandong Guo , Mingming Gong , Tongliang Liu

Articulated object generation has seen increasing advancements, yet existing models often lack the ability to be conditioned on text prompts. To address the significant gap between textual descriptions and 3D articulated object…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Hao Sun , Lei Fan , Donglin Di , Shaohui Liu

Geometric organization of objects into semantically meaningful arrangements pervades the built world. As such, assistive robots operating in warehouses, offices, and homes would greatly benefit from the ability to recognize and rearrange…

Robotics · Computer Science 2021-10-22 Weiyu Liu , Chris Paxton , Tucker Hermans , Dieter Fox