English
Related papers

Related papers: DIPO: Dual-State Images Controlled Articulated Obj…

200 papers

Imitation Learning (IL) is a promising paradigm for learning dynamic manipulation of deformable objects since it does not depend on difficult-to-create accurate simulations of such objects. However, the translation of motions demonstrated…

Robotics · Computer Science 2024-03-20 Eric Hannus , Tran Nguyen Le , David Blanco-Mulero , Ville Kyrki

Deep generative models have been recently extended to synthesizing 3D digital humans. However, previous approaches treat clothed humans as a single chunk of geometry without considering the compositionality of clothing and accessories. As a…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Taeksoo Kim , Shunsuke Saito , Hanbyul Joo

Text-to-3D generation automates 3D content creation from textual descriptions, which offers transformative potential across various fields. However, existing methods often struggle to align generated content with human preferences, limiting…

Computation and Language · Computer Science 2025-02-10 Zhenglin Zhou , Xiaobo Xia , Fan Ma , Hehe Fan , Yi Yang , Tat-Seng Chua

3D Gaussian Splatting, known for enabling high-quality static scene reconstruction with fast rendering, is increasingly being applied to multi-view dynamic scene reconstruction. A common strategy involves learning a deformation field to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Han Jiao , Jiakai Sun , Yexing Xu , Lei Zhao , Wei Xing , Huaizhong Lin

3D Human motion generation is pivotal across film, animation, gaming, and embodied intelligence. Traditional 3D motion synthesis relies on costly motion capture, while recent work shows that 2D videos provide rich, temporally coherent…

Graphics · Computer Science 2026-05-20 Yi-Yang Zhang , Tengjiao Sun , Pengcheng Fang , Deng-Bao Wang , Xiaohao Cai , Min-Ling Zhang , Hansung Kim

This paper introduces MIDI, a novel paradigm for compositional 3D scene generation from a single image. Unlike existing methods that rely on reconstruction or retrieval techniques or recent approaches that employ multi-stage…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Zehuan Huang , Yuan-Chen Guo , Xingqiao An , Yunhan Yang , Yangguang Li , Zi-Xin Zou , Ding Liang , Xihui Liu , Yan-Pei Cao , Lu Sheng

Articulated objects are ubiquitous in daily life. In this paper, we present DexSim2Real$^{2}$, a novel framework for goal-conditioned articulated object manipulation. The core of our framework is constructing an explicit world model of…

Robotics · Computer Science 2025-07-15 Taoran Jiang , Yixuan Guan , Liqian Ma , Jing Xu , Jiaojiao Meng , Weihang Chen , Zecui Zeng , Lusong Li , Dan Wu , Rui Chen

Imitation learning for robotic manipulation has progressed from 2D image policies to 3D representations that explicitly encode geometry. Yet purely geometric policies often lack explicit part-level semantics, which are critical for…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Chongyang Xu , Shen Cheng , Haipeng Li , Haoqiang Fan , Ziliang Feng , Shuaicheng Liu

3D scene graphs have empowered robots with semantic understanding for navigation and planning. However, current functional scene graphs primarily focus on static element detection, lacking the actionable kinematic information required for…

Despite tremendous recent progress in human video generation, generative video diffusion models still struggle to capture the dynamics and physics of human motions faithfully. In this paper, we propose a new framework for human video…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Tao Hu , Varun Jampani

We present Hand ArticuLated Occupancy (HALO), a novel representation of articulated hands that bridges the advantages of 3D keypoints and neural implicit surfaces and can be used in end-to-end trainable architectures. Unlike existing…

Computer Vision and Pattern Recognition · Computer Science 2021-09-24 Korrawe Karunratanakul , Adrian Spurr , Zicong Fan , Otmar Hilliges , Siyu Tang

In this paper, we introduce DINO-X, which is a unified object-centric vision model developed by IDEA Research with the best open-world object detection performance to date. DINO-X employs the same Transformer-based encoder-decoder…

Recent advances in generative models have produced strong results for static 3D shapes, whereas articulated 3D generation remains challenging due to action-dependent deformations and limited datasets. We introduce ArticFlow, a two-stage…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Jiong Lin , Jinchen Ruan , Hod Lipson

Perception of deformable linear objects (DLOs), such as cables, ropes, and wires, is the cornerstone for successful downstream manipulation. Although vision-based methods have been extensively explored, they remain highly vulnerable to…

Robotics · Computer Science 2025-12-22 Kangchen Lv , Mingrui Yu , Shihefeng Wang , Xiangyang Ji , Xiang Li

High-quality human motion data is becoming increasingly important for applications in robotics, simulation, and entertainment. Recent generative models offer a potential data source, enabling human motion synthesis through intuitive inputs…

Learning from demonstrations faces challenges in generalizing beyond the training data and often lacks collision awareness. This paper introduces Lan-o3dp, a language-guided object-centric diffusion policy framework that can adapt to unseen…

Robotics · Computer Science 2025-03-18 Hang Li , Qian Feng , Zhi Zheng , Jianxiang Feng , Zhaopeng Chen , Alois Knoll

Synthesizing whole-body manipulation of articulated objects, including body motion, hand motion, and object motion, is a critical yet challenging task with broad applications in virtual humans and robotics. The core challenges are twofold.…

Graphics · Computer Science 2025-05-28 Huaijin Pi , Zhi Cen , Zhiyang Dou , Taku Komura

Articulated objects exist widely in the real world. However, previous 3D generative methods for unsupervised part decomposition are unsuitable for such objects, because they assume a spatially fixed part location, resulting in inconsistent…

Computer Vision and Pattern Recognition · Computer Science 2021-10-12 Yuki Kawana , Yusuke Mukuta , Tatsuya Harada

Simulating object dynamics from real-world perception shows great promise for digital twins and robotic manipulation but often demands labor-intensive measurements and expertise. We present a fully automated Real2Sim pipeline that generates…

Robotics · Computer Science 2025-04-02 Nicholas Pfaff , Evelyn Fu , Jeremy Binagia , Phillip Isola , Russ Tedrake

Learning bimanual manipulation is challenging due to its high dimensionality and tight coordination required between two arms. Eye-in-hand imitation learning, which uses wrist-mounted cameras, simplifies perception by focusing on…

Robotics · Computer Science 2025-08-19 I-Chun Arthur Liu , Jason Chen , Gaurav Sukhatme , Daniel Seita
‹ Prev 1 8 9 10 Next ›