中文
相关论文

相关论文: UMIGen: A Unified Framework for Egocentric Point C…

200 篇论文

Recently,vision-based robotic manipulation has garnered significant attention and witnessed substantial advancements. 2D image-based and 3D point cloud-based policy learning represent two predominant paradigms in the field, with recent…

机器人学 · 计算机科学 2025-09-23 Run Yu , Yangdi Liu , Wen-Da Wei , Chen Li

We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to achieve unification along three axes: the model, the tasks,…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Chi Zhang , Jiepeng Wang , Youming Wang , Yuanzhi Liang , Xiaoyan Yang , Zuoxin Li , Haibin Huang , Xuelong Li

In this paper, we adopt the Universal Manifold Embedding (UME) framework for the estimation of rigid transformations and extend it, so that it can accommodate scenarios involving partial overlap and differently sampled point clouds. UME is…

计算机视觉与模式识别 · 计算机科学 2024-08-26 Yuval Haitman , Amit Efraim , Joseph M. Francos

In the era of deep learning, data is the critical determining factor in the performance of neural network models. Generating large datasets suffers from various difficulties such as scalability, cost efficiency and photorealism. To avoid…

计算机视觉与模式识别 · 计算机科学 2022-10-04 Chahat Deep Singh , Riya Kumari , Cornelia Fermüller , Nitin J. Sanket , Yiannis Aloimonos

We present Universal Manipulation Interface (UMI) -- a data collection and policy learning framework that allows direct skill transfer from in-the-wild human demonstrations to deployable robot policies. UMI employs hand-held grippers…

机器人学 · 计算机科学 2024-03-07 Cheng Chi , Zhenjia Xu , Chuer Pan , Eric Cousineau , Benjamin Burchfiel , Siyuan Feng , Russ Tedrake , Shuran Song

Large-scale real-world robot data collection is a prerequisite for bringing robots into everyday deployment. However, existing pipelines often rely on specialized handheld devices to bridge the embodiment gap, which not only increases…

机器人学 · 计算机科学 2026-04-10 Yanwen Zou , Chenyang Shi , Wenye Yu , Han Xue , Jun Lv , Ye Pan , Chuan Wen , Cewu Lu

Egocentric human motion generation and forecasting with scene-context is crucial for enhancing AR/VR experiences, improving human-robot interaction, advancing assistive technologies, and enabling adaptive healthcare solutions by accurately…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Chaitanya Patel , Hiroki Nakamura , Yuta Kyuragi , Kazuki Kozuka , Juan Carlos Niebles , Ehsan Adeli

The past few years have witnessed the rapid development of vision-centric 3D perception in autonomous driving. Although the 3D perception models share many structural and conceptual similarities, there still exist gaps in their feature…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Yu Hong , Qian Liu , Huayuan Cheng , Danjiao Ma , Hang Dai , Yu Wang , Guangzhi Cao , Yong Ding

Unified understanding and generation is a highly appealing research direction in multimodal learning. There exist two approaches: one trains a transformer via an auto-regressive paradigm, and the other adopts a two-stage scheme connecting…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Shihao Zhao , Yitong Chen , Zeyinzi Jiang , Bojia Zi , Shaozhe Hao , Yu Liu , Chaojie Mao , Kwan-Yee K. Wong

Robust imitation learning for robot manipulation requires comprehensive 3D perception, yet many existing methods struggle in cluttered environments. Fixed camera view approaches are vulnerable to perspective changes, and 3D point cloud…

机器人学 · 计算机科学 2025-07-08 Daqi Huang , Zhehao Cai , Yuzhi Hao , Zechen Li , Chee-Meng Chew

In recent years, significant progress has been made in both image generation and generated image detection. Despite their rapid, yet largely independent, development, these two fields have evolved distinct architectural paradigms: the…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Yanran Zhang , Wenzhao Zheng , Yifei Li , Bingyao Yu , Yu Zheng , Lei Chen , Jiwen Lu , Jie Zhou

Recent advances in Vision-Language-Action (VLA) and world-model methods have improved generalization in tasks such as robotic manipulation and object interaction. However, Successful execution of such tasks depends on large, costly…

机器人学 · 计算机科学 2026-03-16 Yulu Wu , Jiujun Cheng , Haowen Wang , Dengyang Suo , Pei Ren , Qichao Mao , Shangce Gao , Yakun Huang

This article presents a 3D point cloud map-merging framework for egocentric heterogeneous multi-robot exploration, based on overlap detection and alignment, that is independent of a manual initial guess or prior knowledge of the robots'…

机器人学 · 计算机科学 2023-11-21 Nikolaos Stathoulopoulos , Anton Koval , Ali-akbar Agha-mohammadi , George Nikolakopoulos

The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos…

机器人学 · 计算机科学 2024-11-01 Simar Kareer , Dhruv Patel , Ryan Punamiya , Pranay Mathur , Shuo Cheng , Chen Wang , Judy Hoffman , Danfei Xu

We present Uni$^2$Det, a brand new framework for unified and universal multi-dataset training on 3D detection, enabling robust performance across diverse domains and generalization to unseen domains. Due to substantial disparities in data…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Yubin Wang , Zhikang Zou , Xiaoqing Ye , Xiao Tan , Errui Ding , Cairong Zhao

Mobile imitation learning on portable demonstration interfaces faces two coupled bottlenecks: locomotion-contaminated action labels and inference-induced execution latency on a continuously moving base. Recent wrist-mounted interfaces lower…

机器人学 · 计算机科学 2026-05-21 Haoran Huang , Haonan Dong , Huixu Dong

Utilizing Vision-Language Models (VLMs) for robotic manipulation represents a novel paradigm, aiming to enhance the model's ability to generalize to new objects and instructions. However, due to variations in camera specifications and…

机器人学 · 计算机科学 2024-09-13 Fanfan Liu , Feng Yan , Liming Zheng , Chengjian Feng , Yiyang Huang , Lin Ma

The acquisition of annotated datasets with paired images and segmentation masks is a critical challenge in domains such as medical imaging, remote sensing, and computer vision. Manual annotation demands significant resources, faces ethical…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Rupak Bose , Chinedu Innocent Nwoye , Aditya Bhat , Nicolas Padoy

Recently, robotics has advanced significantly through the integration of larger models and large-scale datasets. However, challenges remain in applying these models to 3D spatial interactions and managing data collection costs. To address…

机器人学 · 计算机科学 2025-11-05 Feng Yan , Fanfan Liu , Liming Zheng , Yufeng Zhong , Yiyang Huang , Zechao Guan , Chengjian Feng , Lin Ma

Human generation has achieved significant progress. Nonetheless, existing methods still struggle to synthesize specific regions such as faces and hands. We argue that the main reason is rooted in the training data. A holistic human dataset…

计算机视觉与模式识别 · 计算机科学 2023-09-26 Jianglin Fu , Shikai Li , Yuming Jiang , Kwan-Yee Lin , Wayne Wu , Ziwei Liu