中文
相关论文

相关论文: Articulate3D: Zero-Shot Text-Driven 3D Object Posi…

200 篇论文

Despite the recent success of multi-view diffusion models for text/image-based 3D asset generation, instruction-based editing of 3D assets lacks surprisingly far behind the quality of generation models. The main reason is that recent…

图形学 · 计算机科学 2025-12-15 Maria Parelli , Michael Oechsle , Michael Niemeyer , Federico Tombari , Andreas Geiger

Effectively manipulating articulated objects in household scenarios is a crucial step toward achieving general embodied artificial intelligence. Mainstream research in 3D vision has primarily focused on manipulation through depth perception…

机器人学 · 计算机科学 2025-03-24 Wenbo Cui , Chengyang Zhao , Songlin Wei , Jiazhao Zhang , Haoran Geng , Yaran Chen , Haoran Li , He Wang

Estimating the 3D hand articulation from a single color image is an important problem with applications in Augmented Reality (AR), Virtual Reality (VR), Human-Computer Interaction (HCI), and robotics. Apart from the absence of depth…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Christos Pantazopoulos , Spyridon Thermos , Gerasimos Potamianos

While much progress has been made on the task of 3D point cloud registration, there still exists no learning-based method able to estimate the 6D pose of an object observed by a 2.5D sensor in a scene. The challenges of this scenario…

计算机视觉与模式识别 · 计算机科学 2020-11-24 Zheng Dang , Fei Wang , Mathieu Salzmann

Generative methods for 3D assets have recently achieved remarkable progress, yet providing intuitive and precise control over the object geometry remains a key challenge. Existing approaches predominantly rely on text or image prompts,…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Elisabetta Fedele , Francis Engelmann , Ian Huang , Or Litany , Marc Pollefeys , Leonidas Guibas

Articulated object generation has seen increasing advancements, yet existing models often lack the ability to be conditioned on text prompts. To address the significant gap between textual descriptions and 3D articulated object…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Hao Sun , Lei Fan , Donglin Di , Shaohui Liu

We introduce Dream2Real, a robotics framework which integrates vision-language models (VLMs) trained on 2D data into a 3D object rearrangement pipeline. This is achieved by the robot autonomously constructing a 3D representation of the…

机器人学 · 计算机科学 2024-07-31 Ivan Kapelyukh , Yifei Ren , Ignacio Alzugaray , Edward Johns

Articulated objects are ubiquitous in daily life. Our goal is to achieve a high-quality reconstruction, segmentation of independent moving parts, and analysis of articulation. Recent methods analyse two different articulation states and…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Hao Ai , Wenjie Chang , Jianbo Jiao , Ales Leonardis , Ofek Eyal

Rigged objects are commonly used in artist pipelines, as they can flexibly adapt to different scenes and postures. However, articulating the rigs into realistic affordance-aware postures (e.g., following the context, respecting the physics…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Yu-Chu Yu , Chieh Hubert Lin , Hsin-Ying Lee , Chaoyang Wang , Yu-Chiang Frank Wang , Ming-Hsuan Yang

We study the problem of learning to estimate the 3D object pose from a few labelled examples and a collection of unlabelled data. Our main contribution is a learning framework, neural view synthesis and matching, that can transfer the 3D…

计算机视觉与模式识别 · 计算机科学 2021-10-28 Angtian Wang , Shenxiao Mei , Alan Yuille , Adam Kortylewski

Human is able to conduct 3D recognition by a limited number of haptic contacts between the target object and his/her fingers without seeing the object. This capability is defined as `haptic glance' in cognitive neuroscience. Most of the…

人工智能 · 计算机科学 2021-02-16 Kevin Riou , Suiyi Ling , Guillaume Gallot , Patrick Le Callet

Current state-of-the-art methods cast monocular 3D human pose estimation as a learning problem by training neural networks on large data sets of images and corresponding skeleton poses. In contrast, we propose an approach that can exploit…

计算机视觉与模式识别 · 计算机科学 2020-10-14 Simon Jenni , Paolo Favaro

We present a novel approach to the generation of static and articulated 3D assets that has a 3D autodecoder at its core. The 3D autodecoder framework embeds properties learned from the target dataset in the latent space, which can then be…

计算机视觉与模式识别 · 计算机科学 2023-07-12 Evangelos Ntavelis , Aliaksandr Siarohin , Kyle Olszewski , Chaoyang Wang , Luc Van Gool , Sergey Tulyakov

Estimating the 3D pose of desktop objects is crucial for applications such as robotic manipulation. Many existing approaches to this problem require a depth map of the object for both training and prediction, which restricts them to opaque,…

计算机视觉与模式识别 · 计算机科学 2020-05-20 Xingyu Liu , Rico Jonschkowski , Anelia Angelova , Kurt Konolige

We present Point2Pose, a model-free method for causal 6D pose tracking of multiple rigid objects from monocular RGB-D video. Initialized only from sparse image points on the objects to be tracked, our approach tracks multiple unseen objects…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Tzu-Yuan Lin , Ho Jae Lee , Kevin Doherty , Yonghyeon Lee , Sangbae Kim

Image Matching is a core component of all best-performing algorithms and pipelines in 3D vision. Yet despite matching being fundamentally a 3D problem, intrinsically linked to camera pose and scene geometry, it is typically treated as a 2D…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Vincent Leroy , Yohann Cabon , Jérôme Revaud

3D Visual Grounding (3DVG) is an essential capability for embodied AI, requiring agents to localize objects in 3D scenes based on natural language descriptions. Recent zero-shot methods leverage 2D vision-language models (LVLMs). However,…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Cuong Huynh , Maxim Popov , Denis Gridusov , Sergey Kolyubin

Driven by powerful image diffusion models, recent research has achieved the automatic creation of 3D objects from textual or visual guidance. By performing score distillation sampling (SDS) iteratively across different views, these methods…

计算机视觉与模式识别 · 计算机科学 2024-05-17 Zeyu Li , Ruitong Gan , Chuanchen Luo , Yuxi Wang , Jiaheng Liu , Ziwei Zhu Man Zhang , Qing Li , Xucheng Yin , Zhaoxiang Zhang , Junran Peng

State-of-the-art computer vision algorithms often achieve efficiency by making discrete choices about which hypotheses to explore next. This allows allocation of computational resources to promising candidates, however, such decisions are…

计算机视觉与模式识别 · 计算机科学 2017-04-12 Alexander Krull , Eric Brachmann , Sebastian Nowozin , Frank Michel , Jamie Shotton , Carsten Rother

In this paper, we introduce a new task: Zero-Shot 3D Reasoning Segmentation for parts searching and localization for objects, which is a new paradigm to 3D segmentation that transcends limitations for previous category-specific 3D semantic…

计算机视觉与模式识别 · 计算机科学 2024-05-30 Tianrun Chen , Chunan Yu , Jing Li , Jianqi Zhang , Lanyun Zhu , Deyi Ji , Yong Zhang , Ying Zang , Zejian Li , Lingyun Sun