中文
相关论文

相关论文: SPLIT: SE(3)-diffusion via Local Geometry-based Sc…

200 篇论文

We present iFusion, a novel 3D object reconstruction framework that requires only two views with unknown camera poses. While single-view reconstruction yields visually appealing results, it can deviate significantly from the actual object,…

计算机视觉与模式识别 · 计算机科学 2023-12-29 Chin-Hsuan Wu , Yen-Chun Chen , Bolivar Solarte , Lu Yuan , Min Sun

Pose prediction is to predict future poses given a window of previous poses. In this paper, we propose a new problem that predicts poses using 3D joint coordinate sequences. Different from the traditional pose prediction based on Mocap…

计算机视觉与模式识别 · 计算机科学 2019-09-05 Xiaoli Liu , Jianqin Yin , Huaping Liu , Yilong Yin

Hand pose estimation from a single image has many applications. However, approaches to full 3D body pose estimation are typically trained on day-to-day activities or actions. As such, detailed hand-to-hand interactions are poorly…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Maksym Ivashechkin , Oscar Mendez , Richard Bowden

Understanding visual scenes is fundamental to human intelligence. While discriminative models have significantly advanced computer vision, they often struggle with compositional understanding. In contrast, recent generative text-to-image…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Yujin Jeong , Arnas Uselis , Seong Joon Oh , Anna Rohrbach

6D pose estimation of rigid objects from RGB-D images is crucial for object grasping and manipulation in robotics. Although RGB channels and the depth (D) channel are often complementary, providing respectively the appearance and geometry…

计算机视觉与模式识别 · 计算机科学 2022-08-18 Haoran Pan , Jun Zhou , Yuanpeng Liu , Xuequan Lu , Weiming Wang , Xuefeng Yan , Mingqiang Wei

We introduce a novel, training-free system for reconstructing, understanding, and rendering 3D indoor scenes from a sparse set of unposed RGB images. Unlike traditional radiance field approaches that require dense views and per-scene…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Jiatong Xia , Lingqiao Liu

The remarkable achievements of both generative models of 2D images and neural field representations for 3D scenes present a compelling opportunity to integrate the strengths of both approaches. In this work, we propose a methodology that…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Azmi Haider , Dan Rosenbaum

Shape assembly aims to reassemble parts (or fragments) into a complete object, which is a common task in our daily life. Different from the semantic part assembly (e.g., assembling a chair's semantic parts like legs into a whole chair),…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Ruihai Wu , Chenrui Tie , Yushi Du , Yan Zhao , Hao Dong

Significant strides have been made using large vision-language models, like Stable Diffusion (SD), for a variety of downstream tasks, including image editing, image correspondence, and 3D shape generation. Inspired by these advancements, we…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Aliasghar Khani , Saeid Asgari Taghanaki , Aditya Sanghi , Ali Mahdavi Amiri , Ghassan Hamarneh

Robotic manipulation in unstructured environments requires the generation of robust and long-horizon trajectory-level policy with conditions of perceptual observations and benefits from the advantages of SE(3)-equivariant diffusion models…

机器人学 · 计算机科学 2025-09-30 Zhitao Wang , Yanke Wang , Jiangtao Wen , Roberto Horowitz , Yuxing Han

Score-based diffusion models learn to reverse a stochastic differential equation that maps data to noise. However, for complex tasks, numerical error can compound and result in highly unnatural samples. Previous work mitigates this drift…

机器学习 · 统计学 2023-06-12 Aaron Lou , Stefano Ermon

Precise manipulation that is generalizable across scenes and objects remains a persistent challenge in robotics. Current approaches for this task heavily depend on having a significant number of training instances to handle objects with…

机器人学 · 计算机科学 2024-12-30 Nikolaos Tsagkas , Jack Rome , Subramanian Ramamoorthy , Oisin Mac Aodha , Chris Xiaoxuan Lu

Annotating 3D LiDAR point clouds for perception tasks is fundamental for many applications e.g., autonomous driving, yet it still remains notoriously labor-intensive. Pretraining-finetuning approach can alleviate the labeling burden by…

计算机视觉与模式识别 · 计算机科学 2025-07-17 Xiangchao Yan , Runjian Chen , Bo Zhang , Hancheng Ye , Renqiu Xia , Jiakang Yuan , Hongbin Zhou , Xinyu Cai , Botian Shi , Wenqi Shao , Ping Luo , Yu Qiao , Tao Chen , Junchi Yan

Diffusion models generate images with an unprecedented level of quality, but how can we freely rearrange image layouts? Recent works generate controllable scenes via learning spatially disentangled latent codes, but these methods do not…

计算机视觉与模式识别 · 计算机科学 2024-04-11 Jiawei Ren , Mengmeng Xu , Jui-Chieh Wu , Ziwei Liu , Tao Xiang , Antoine Toisoul

Semantic segmentation and activity classification are key components to creating intelligent surgical systems able to understand and assist clinical workflow. In the Operating Room, semantic segmentation is at the core of creating robots…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Idris Hamoud , Alexandros Karargyris , Aidean Sharghi , Omid Mohareri , Nicolas Padoy

3D object detection and pose estimation has been studied extensively in recent decades for its potential applications in robotics. However, there still remains challenges when we aim at detecting multiple objects while retaining low false…

机器人学 · 计算机科学 2017-03-14 Ruotao He , Juan Rojas , Yisheng Guan

We introduce S2C-3D, a novel sparse-view 3D reconstruction framework for high-fidelity and complete scene reconstruction from as few as six to eight images. Our framework features three components: a specialized diffusion model for…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Yiyang Shen , Yin Yang , Kun Zhou , Tianjia Shao

Scene understanding remains a significant challenge in the computer vision community. The visual psychophysics literature has demonstrated the importance of interdependence among parts of the scene. Yet, the majority of methods in computer…

计算机视觉与模式识别 · 计算机科学 2011-08-23 Jason J. Corso

Real-time holistic scene understanding would allow machines to interpret their surrounding in a much more detailed manner than is currently possible. While panoptic image segmentation methods have brought image segmentation closer to this…

计算机视觉与模式识别 · 计算机科学 2022-04-04 Leevi Raivio , Esa Rahtu

Large pre-trained models have had a significant impact on computer vision by enabling multi-modal learning, where the CLIP model has achieved impressive results in image classification, object detection, and semantic segmentation. However,…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Sitian Shen , Zilin Zhu , Linqian Fan , Harry Zhang , Xinxiao Wu