中文
相关论文

相关论文: PartRM: Modeling Part-Level Dynamics with Large Cr…

200 篇论文

Can we scale 4D pretraining to learn general space-time representations that reconstruct an object from a few views at some times to any view at any time? We provide an affirmative answer with 4D-LRM, the first large-scale 4D reconstruction…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Ziqiao Ma , Xuweiyi Chen , Shoubin Yu , Sai Bi , Kai Zhang , Chen Ziwen , Sihan Xu , Jianing Yang , Zexiang Xu , Kalyan Sunkavalli , Mohit Bansal , Joyce Chai , Hao Tan

We introduce Puppet-Master, an interactive video generator that captures the internal, part-level motion of objects, serving as a proxy for modeling object dynamics universally. Given an image of an object and a set of "drags" specifying…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Ruining Li , Chuanxia Zheng , Christian Rupprecht , Andrea Vedaldi

Real-world objects are composed of distinctive, object-specific parts. Identifying these parts is key to performing fine-grained, compositional reasoning-yet, large multimodal models (LMMs) struggle to perform this seemingly straightforward…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Ansel Blume , Jeonghwan Kim , Hyeonjeong Ha , Elen Chatikyan , Xiaomeng Jin , Khanh Duy Nguyen , Nanyun Peng , Kai-Wei Chang , Derek Hoiem , Heng Ji

Large Reconstruction Models (LRMs) have recently become a popular method for creating 3D foundational models. Training 3D reconstruction models with 2D visual data traditionally requires prior knowledge of camera poses for the training…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Shiu-hong Kao , Xiao Li , Jinglu Wang , Yang Li , Chi-Keung Tang , Yu-Wing Tai , Yan Lu

Modeling 3D articulated objects with realistic geometry, textures, and kinematics is essential for a wide range of applications. However, existing optimization-based reconstruction methods often require dense multi-view inputs and expensive…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Sylvia Yuan , Ruoxi Shi , Xinyue Wei , Xiaoshuai Zhang , Hao Su , Minghua Liu

Fine-grained robot manipulation, such as lifting and rotating a bottle to display the label on the cap, requires robust reasoning about object parts and their relationships with intended tasks. Despite recent advances in training…

机器人学 · 计算机科学 2025-06-18 Yifan Yin , Zhengtao Han , Shivam Aarya , Jianxin Wang , Shuhang Xu , Jiawei Peng , Angtian Wang , Alan Yuille , Tianmin Shu

Existing methods for reconstructing interactive scenes primarily focus on replacing reconstructed objects with CAD models retrieved from a limited database, resulting in significant discrepancies between the reconstructed and observed…

机器人学 · 计算机科学 2023-08-02 Zeyu Zhang , Lexing Zhang , Zaijin Wang , Ziyuan Jiao , Muzhi Han , Yixin Zhu , Song-Chun Zhu , Hangxin Liu

DUSt3R has recently shown that one can reduce many tasks in multi-view geometry, including estimating camera intrinsics and extrinsics, reconstructing the scene in 3D, and establishing image correspondences, to the prediction of a pair of…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Edgar Sucar , Zihang Lai , Eldar Insafutdinov , Andrea Vedaldi

The default strategy for training single-view Large Reconstruction Models (LRMs) follows the fully supervised route using large-scale datasets of synthetic 3D assets or multi-view captures. Although these resources simplify the training…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Hanwen Jiang , Qixing Huang , Georgios Pavlakos

Understanding objects at the level of their constituent parts is fundamental to advancing computer vision, graphics, and robotics. While datasets like PartNet have driven progress in 3D part understanding, their reliance on untextured…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Penghao Wang , Yiyang He , Xin Lv , Yukai Zhou , Lan Xu , Jingyi Yu , Jiayuan Gu

Part-level representations are essential for robust person re-identification. However, common errors that arise during pedestrian detection frequently result in severe misalignment problems for body parts, which degrade the quality of part…

计算机视觉与模式识别 · 计算机科学 2019-12-16 Kan Wang , Changxing Ding , Stephen J. Maybank , Dacheng Tao

We present Large Inverse Rendering Model (LIRM), a transformer architecture that jointly reconstructs high-quality shape, materials, and radiance fields with view-dependent effects in less than a second. Our model builds upon the recent…

3D human reconstruction and animation are long-standing topics in computer graphics and vision. However, existing methods typically rely on sophisticated dense-view capture and/or time-consuming per-subject optimization procedures. To…

图形学 · 计算机科学 2025-06-04 Zhiyuan Yu , Zhe Li , Hujun Bao , Can Yang , Xiaowei Zhou

Robust imitation learning for robot manipulation requires comprehensive 3D perception, yet many existing methods struggle in cluttered environments. Fixed camera view approaches are vulnerable to perspective changes, and 3D point cloud…

机器人学 · 计算机科学 2025-07-08 Daqi Huang , Zhehao Cai , Yuzhi Hao , Zechen Li , Chee-Meng Chew

Articulated 3D objects are critical for embodied AI, robotics, and interactive scene understanding, yet creating simulation-ready assets remains labor-intensive and requires expert modeling of part hierarchies and motion structures. We…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Yumeng He , Ying Jiang , Jiayin Lu , Yin Yang , Chenfanfu Jiang

Humans can learn to manipulate new objects by simply watching others; providing robots with the ability to learn from such demonstrations would enable a natural interface specifying new behaviors. This work develops Robot See Robot Do…

机器人学 · 计算机科学 2024-09-27 Justin Kerr , Chung Min Kim , Mingxuan Wu , Brent Yi , Qianqian Wang , Ken Goldberg , Angjoo Kanazawa

Precise 6D pose estimation of rigid objects from RGB images is a critical but challenging task in robotics, augmented reality and human-computer interaction. To address this problem, we propose DeepRM, a novel recurrent network architecture…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Alexander Avery , Andreas Savakis

We propose a Pose-Free Large Reconstruction Model (PF-LRM) for reconstructing a 3D object from a few unposed images even with little visual overlap, while simultaneously estimating the relative camera poses in ~1.3 seconds on a single A100…

计算机视觉与模式识别 · 计算机科学 2023-11-27 Peng Wang , Hao Tan , Sai Bi , Yinghao Xu , Fujun Luan , Kalyan Sunkavalli , Wenping Wang , Zexiang Xu , Kai Zhang

Existing text-driven 3D human motion editing methods have demonstrated significant progress, but are still difficult to precisely control over detailed, part-specific motions due to their global modeling nature. In this paper, we propose…

图形学 · 计算机科学 2026-01-01 Yujie Yang , Zhichao Zhang , Jiazhou Chen , Zichao Wu

Foundation models pre-trained on massive unlabeled datasets have revolutionized natural language and computer vision, exhibiting remarkable generalization capabilities, thus highlighting the importance of pre-training. Yet, efforts in…

机器人学 · 计算机科学 2025-05-20 Dantong Niu , Yuvan Sharma , Haoru Xue , Giscard Biamby , Junyi Zhang , Ziteng Ji , Trevor Darrell , Roei Herzig
‹ 上一页 1 2 3 10 下一页 ›