中文
相关论文

相关论文: EnerVerse: Envisioning Embodied Future Space for R…

200 篇论文

This paper tackles the challenge of robust reconstruction, i.e., the task of reconstructing a 3D scene from a set of inconsistent multi-view images. Some recent works have attempted to simultaneously remove image inconsistencies and perform…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Jin Cao , Hongrui Wu , Ziyong Feng , Hujun Bao , Xiaowei Zhou , Sida Peng

Enabling robots to execute long-horizon manipulation tasks from free-form language instructions remains a fundamental challenge in embodied AI. While vision-language models (VLMs) have shown promise as high-level planners, their deployment…

机器人学 · 计算机科学 2025-10-01 Zitong Bo , Yue Hu , Jinming Ma , Mingliang Zhou , Junhui Yin , Yachen Kang , Yuqi Liu , Tong Wu , Diyun Xiang , Hao Chen

Egocentric human videos provide a scalable source of manipulation demonstrations; however, deploying them on robots requires active viewpoint control to maintain task-critical visibility, which human viewpoint imitation often fails to…

机器人学 · 计算机科学 2026-02-27 Daesol Cho , Youngseok Jang , Danfei Xu , Sehoon Ha

End-effector based assistive robots face persistent challenges in generating smooth and robust trajectories when controlled by human's noisy and unreliable biosignals such as muscle activities and brainwaves. The produced endpoint…

机器人学 · 计算机科学 2025-06-12 Ali Rabiee , Sima Ghafoori , Xiangyu Bai , Sarah Ostadabbas , Reza Abiri

Current machine learning techniques proposed to automatically discover a robot kinematics usually rely on a priori information about the robot's structure, sensors properties or end-effector position. This paper proposes a method to…

机器学习 · 计算机科学 2018-10-05 Alban Laflaquière , Alexander V. Terekhov , Bruno Gas , J. Kevin O'Regan

This paper introduces a pioneering 3D volumetric encoder designed for text-to-3D generation. To scale up the training data for the diffusion model, a lightweight network is developed to efficiently acquire feature volumes from multi-view…

计算机视觉与模式识别 · 计算机科学 2024-08-14 Zhicong Tang , Shuyang Gu , Chunyu Wang , Ting Zhang , Jianmin Bao , Dong Chen , Baining Guo

Virtual Reality, and Extended Reality in general, connect the physical body with the virtual world. Movement of our body translates to interactions with this virtual world. Only by moving our head will we see a different perspective. By…

多媒体 · 计算机科学 2022-12-09 Glenn Van Wallendael , Lucas Liegeois , Julie Artois , Peter Lambert

Ergodic control synthesizes optimal coverage behaviors over spatial distributions for nonlinear systems. However, existing formulations model the robot as a non-volumetric point, whereas in practice a robot interacts with the environment…

机器人学 · 计算机科学 2026-04-14 Jueun Kwon , Max M. Sun , Todd Murphey

We introduce DreamGen, a simple yet highly effective 4-stage pipeline for training robot policies that generalize across behaviors and environments through neural trajectories - synthetic robot data generated from video world models.…

The metaverse is expected to provide immersive entertainment, education, and business applications. However, virtual reality (VR) transmission over wireless networks is data- and computation-intensive, making it critical to introduce novel…

图像与视频处理 · 电气工程与系统科学 2024-01-03 Yiyu Guo , Zhijin Qin , Xiaoming Tao , Geoffrey Ye Li

We propose X-WAM, a Unified 4D World Model that unifies real-time robotic action execution and high-fidelity 4D world synthesis (video + 3D reconstruction) in a single framework, addressing the critical limitations of prior unified world…

机器人学 · 计算机科学 2026-05-08 Jun Guo , Qiwei Li , Peiyan Li , Zilong Chen , Nan Sun , Yifei Su , Heyun Wang , Yuan Zhang , Xinghang Li , Huaping Liu

Visual 3D motion estimation aims to infer the motion of 2D pixels in 3D space based on visual cues. The key challenge arises from depth variation induced spatio-temporal motion inconsistencies, disrupting the assumptions of local spatial or…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Zengyu Wan , Wei Zhai , Yang Cao , Zhengjun Zha

Video-based world models hold significant potential for generating high-quality embodied manipulation data. However, current video generation methods struggle to achieve stable long-horizon generation: classical diffusion-based approaches…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Yu Shang , Lei Jin , Yiding Ma , Xin Zhang , Chen Gao , Wei Wu , Yong Li

Sparse RGBD scene completion is a challenging task especially when considering consistent textures and geometries throughout the entire scene. Different from existing solutions that rely on human-designed text prompts or predefined camera…

计算机视觉与模式识别 · 计算机科学 2024-08-05 Ming-Feng Li , Yueh-Feng Ku , Hong-Xuan Yen , Chi Liu , Yu-Lun Liu , Albert Y. C. Chen , Cheng-Hao Kuo , Min Sun

Motion-controllable video generation is crucial for egocentric applications in virtual reality and embodied AI. However, existing methods often struggle to achieve 3D-consistent fine-grained hand articulation. By adopting on 2D trajectories…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Chenyangguang Zhang , Botao Ye , Boqi Chen , Alexandros Delitzas , Fangjinhua Wang , Marc Pollefeys , Xi Wang

World models are emerging as a foundational paradigm for scalable, data-efficient embodied AI. In this work, we present GigaWorld-0, a unified world model framework designed explicitly as a data engine for Vision-Language-Action (VLA)…

3D-aware visual pretraining has proven effective in improving the performance of downstream robotic manipulation tasks. However, existing methods are constrained to Euclidean embedding spaces, whose flat geometry limits their ability to…

机器人学 · 计算机科学 2026-03-13 Jin Yang , Ping Wei , Yixin Chen , Nanning Zheng

Functional grasping with dexterous robotic hands is a key capability for enabling tool use and complex manipulation, yet progress has been constrained by two persistent bottlenecks: the scarcity of large-scale datasets and the absence of…

机器人学 · 计算机科学 2026-01-09 Xingyi He , Adhitya Polavaram , Yunhao Cao , Om Deshmukh , Tianrui Wang , Xiaowei Zhou , Kuan Fang

Optimizing and refining action execution through exploration and interaction is a promising way for robotic manipulation. However, practical approaches to interaction-driven robotic learning are still underexplored, particularly for…

机器人学 · 计算机科学 2025-09-24 Yibo Peng , Jiahao Yang , Shenhao Yan , Ziyu Huang , Shuang Li , Shuguang Cui , Yiming Zhao , Yatong Han

For embodied agents to infer representations of the underlying 3D physical world they inhabit, they should efficiently combine multisensory cues from numerous trials, e.g., by looking at and touching objects. Despite its importance,…

机器学习 · 计算机科学 2019-11-11 Jae Hyun Lim , Pedro O. Pinheiro , Negar Rostamzadeh , Christopher Pal , Sungjin Ahn