中文
相关论文

相关论文: AGILE: Hand-Object Interaction Reconstruction from…

200 篇论文

We introduce Vid-CamEdit, a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view…

We present Real2Code, a novel approach to reconstructing articulated objects via code generation. Given visual observations of an object, we first reconstruct its part geometry using an image segmentation model and a shape completion model.…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Zhao Mandi , Yijia Weng , Dominik Bauer , Shuran Song

To catch a thrown object, a robot must be able to perceive the object's motion and generate control actions in a timely manner. Rather than explicitly estimating the object's 3D position, this work focuses on a novel approach that…

机器人学 · 计算机科学 2026-02-27 Seongyong Kim , Junhyeon Cho , Kang-Won Lee , Soo-Chul Lim

We formulate the motor system of an interactive avatar as a generative motion model that can drive the body to move through 3D space in a perpetual, realistic, controllable, and responsive manner. Although human motion generation has been…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Yan Zhang , Yao Feng , Alpár Cseke , Nitin Saini , Nathan Bajandas , Nicolas Heron , Michael J. Black

Whole-body Humanoid-Object Interaction (HOI) is bottlenecked by the scarcity of high-fidelity 3D data. While video generative priors offer a promising alternative, existing methods suffer from \textit{Representation Misalignment} due to…

Large-scale articulated objects with high quality are desperately needed for multiple tasks related to embodied AI. Most existing methods for creating articulated objects are either data-driven or simulation based, which are limited by the…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Xinyu Lian , Zichao Yu , Ruiming Liang , Yitong Wang , Li Ray Luo , Kaixu Chen , Yuanzhen Zhou , Qihong Tang , Xudong Xu , Zhaoyang Lyu , Bo Dai , Jiangmiao Pang

Exploration is essential for general-purpose robotic learning, especially in open-ended environments where dense rewards, explicit goals, or task-specific supervision are scarce. Vision-language models (VLMs), with their semantic reasoning…

机器人学 · 计算机科学 2025-09-12 Seungjae Lee , Daniel Ekpo , Haowen Liu , Furong Huang , Abhinav Shrivastava , Jia-Bin Huang

Visual actionable affordance has emerged as a transformative approach in robotics, focusing on perceiving interaction areas prior to manipulation. Traditional methods rely on pixel sampling to identify successful interaction samples or…

机器人学 · 计算机科学 2025-10-10 Taewhan Kim , Hojin Bae , Zeming Li , Xiaoqi Li , Iaroslav Ponomarenko , Ruihai Wu , Hao Dong

Scaling Vision-Language-Action (VLA) models requires massive datasets that are both semantically coherent and physically feasible. However, existing scene generation methods often lack context-awareness, making it difficult to synthesize…

机器人学 · 计算机科学 2026-04-13 Yaru Liu , Ao-bo Wang , Nanyang Ye

Synthetic data generated by video generative models has shown promise for robot learning as a scalable pipeline, but it often suffers from inconsistent action quality due to imperfectly generated videos. Recently, vision-language models…

机器人学 · 计算机科学 2026-02-24 Seungku Kim , Suhyeok Jang , Byungjun Yoon , Dongyoung Kim , John Won , Jinwoo Shin

Humans intuitively understand that inanimate objects do not move by themselves, but that state changes are typically caused by human manipulation (e.g., the opening of a book). This is not yet the case for machines. In part this is because…

计算机视觉与模式识别 · 计算机科学 2023-04-25 Zicong Fan , Omid Taheri , Dimitrios Tzionas , Muhammed Kocabas , Manuel Kaufmann , Michael J. Black , Otmar Hilliges

Being able to reproduce physical phenomena ranging from light interaction to contact mechanics, simulators are becoming increasingly useful in more and more application domains where real-world interaction or labeled data are difficult to…

机器人学 · 计算机科学 2022-09-13 Eric Heiden , Ziang Liu , Vibhav Vineet , Erwin Coumans , Gaurav S. Sukhatme

The landscape of video generation is shifting, from a focus on generating visually appealing clips to building virtual environments that support interaction and maintain physical plausibility. These developments point toward the emergence…

人工智能 · 计算机科学 2026-02-09 Jingtong Yue , Ziqi Huang , Zhaoxi Chen , Xintao Wang , Pengfei Wan , Ziwei Liu

Synthesizing realistic human-object interactions (HOI) in video is challenging due to the complex, instance-specific interaction dynamics of both humans and objects. Incorporating controllability in video generation further adds to the…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Wanyue Zhang , Lin Geng Foo , Thabo Beeler , Rishabh Dabral , Christian Theobalt

We aim to tackle the interesting yet challenging problem of generating videos of diverse and natural human motions from prescribed action categories. The key issue lies in the ability to synthesize multiple distinct motion sequences that…

计算机视觉与模式识别 · 计算机科学 2021-12-21 Chuan Guo , Xinxin Zuo , Sen Wang , Xinshuang Liu , Shihao Zou , Minglun Gong , Li Cheng

Yume aims to use images, text, or videos to create an interactive, realistic, and dynamic world, which allows exploration and control using peripheral devices or neural signals. In this report, we present a preview version of \method, which…

计算机视觉与模式识别 · 计算机科学 2025-07-24 Xiaofeng Mao , Shaoheng Lin , Zhen Li , Chuanhao Li , Wenshuo Peng , Tong He , Jiangmiao Pang , Mingmin Chi , Yu Qiao , Kaipeng Zhang

Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint under limited multi-view data. Existing methods often address these factors separately, resulting in limited…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Zhengwentai Sun , Keru Zheng , Chenghong Li , Hongjie Liao , Xihe Yang , Heyuan Li , Yihao Zhi , Shuliang Ning , Shuguang Cui , Xiaoguang Han

Vision-language-action (VLA) models provide a promising paradigm for scalable robotic manipulation, yet their reliance on success-only behavioral cloning leaves them brittle; lacking corrective training signals, minor execution errors…

Agentic task-solving with Large Language Models (LLMs) requires multi-turn, multi-step interactions, often involving complex function calls and dynamic user-agent exchanges. Existing simulation-based data generation methods for such…

计算与语言 · 计算机科学 2026-02-16 Xingshan Zeng , Weiwen Liu , Lingzhi Wang , Liangyou Li , Fei Mi , Yasheng Wang , Lifeng Shang , Xin Jiang , Qun Liu

Animatable 3D assets, defined as geometry equipped with an articulated skeleton and skinning weights, are fundamental to interactive graphics, embodied agents, and animation production. While recent 3D generative models can synthesize…