中文
相关论文

相关论文: AGILE: Hand-Object Interaction Reconstruction from…

200 篇论文

Real-world image restoration (IR) is inherently complex and often requires combining multiple specialized models to address diverse degradations. Inspired by human problem-solving, we propose AgenticIR, an agentic system that mimics the…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Kaiwen Zhu , Jinjin Gu , Zhiyuan You , Yu Qiao , Chao Dong

This work make the first attempt to generate articulated human motion sequence from a single image. On the one hand, we utilize paired inputs including human skeleton information as motion embedding and a single human image as appearance…

计算机视觉与模式识别 · 计算机科学 2017-09-15 Yichao Yan , Jingwei Xu , Bingbing Ni , Xiaokang Yang

Generative search engines represent a transition from traditional ranking-based retrieval to Large Language Model (LLM)-based synthesis, transforming optimization goals from ranking prominence towards content inclusion. Generative Engine…

人工智能 · 计算机科学 2026-03-24 Jiaqi Yuan , Jialu Wang , Zihan Wang , Qingyun Sun , Ruijie Wang , Jianxin Li

Recent breakthroughs in generative simulation have harnessed Large Language Models (LLMs) to generate diverse robotic task curricula, yet these open-loop paradigms frequently produce linguistically coherent but physically infeasible goals,…

机器人学 · 计算机科学 2026-03-03 Bingchuan Wei , Bingqi Huang , Jingheng Ma , Zeyu zhang , Sen Cui

Human-object-scene interactions (HOSI) generation has broad applications in embodied AI, simulation, and animation. Unlike human-object interaction (HOI) and human-scene interaction (HSI), HOSI generation requires reasoning over dynamic…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Yude Zou , Junji Gong , Xing Gao , Zixuan Li , Tianxing Chen , Guanjie Zheng

Recent advances in 3D human-aware generation have made significant progress. However, existing methods still struggle with generating novel Human Object Interaction (HOI) from text, particularly for open-set objects. We identify three main…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Jinlu Zhang , Yixin Chen , Zan Wang , Jie Yang , Yizhou Wang , Siyuan Huang

Current digital human studies focusing on lip-syncing and body movement are no longer sufficient to meet the growing industrial demand, while human video generation techniques that support interacting with real-world environments (e.g.,…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Yingying Fan , Quanwei Yang , Kaisiyuan Wang , Hang Zhou , Yingying Li , Haocheng Feng , Errui Ding , Yu Wu , Jingdong Wang

Vision-Language-Action (VLA) models have emerged as a promising paradigm for robotic manipulation by leveraging pre-trained vision-language representations. However, current VLA training methods suffer from two critical limitations: poor…

机器人学 · 计算机科学 2026-05-25 Ruofan Jin , Zaixi Zhang

We introduce CRISP, a method that recovers simulatable human motion and scene geometry from monocular video. Prior work on joint human-scene reconstruction relies on data-driven priors and joint optimization with no physics in the loop, or…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Zihan Wang , Jiashun Wang , Jeff Tan , Yiwen Zhao , Jessica Hodgins , Shubham Tulsiani , Deva Ramanan

Recent work has demonstrated the ability of deep reinforcement learning (RL) algorithms to learn complex robotic behaviours in simulation, including in the domain of multi-fingered manipulation. However, such models can be challenging to…

Equipping humanoid robots with versatile interaction skills typically requires either extensive policy training or explicit human-to-robot motion retargeting. However, learning-based policies face prohibitive data collection costs.…

3D scene graphs have empowered robots with semantic understanding for navigation and planning. However, current functional scene graphs primarily focus on static element detection, lacking the actionable kinematic information required for…

机器人学 · 计算机科学 2026-03-24 Qiuyi Gu , Yuze Sheng , Jincheng Yu , Jiahao Tang , Xiaolong Shan , Zhaoyang Shen , Tinghao Yi , Xiaodan Liang , Xinlei Chen , Yu Wang

Visual affordance grounding aims to segment all possible interaction regions between people and objects from an image/video, which is beneficial for many applications, such as robot grasping and action recognition. However, existing methods…

计算机视觉与模式识别 · 计算机科学 2021-08-13 Hongchen Luo , Wei Zhai , Jing Zhang , Yang Cao , Dacheng Tao

Over recent years, diffusion models have facilitated significant advancements in video generation. Yet, the creation of face-related videos still confronts issues such as low facial fidelity, lack of frame consistency, limited editability…

计算机视觉与模式识别 · 计算机科学 2023-12-22 Linze Li , Sunqi Fan , Hengjun Pu , Zhaodong Bing , Yao Tang , Tianzhu Ye , Tong Yang , Liangyu Chen , Jiajun Liang

Large-scale video generative models can synthesize diverse and realistic visual content for dynamic world creation, but they often lack element-wise controllability, hindering their use in editing scenes and training embodied AI agents. We…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Sicheng Mo , Ziyang Leng , Leon Liu , Weizhen Wang , Honglin He , Bolei Zhou

Recent advances in world models have greatly enhanced interactive environment simulation. Existing methods mainly fall into two categories: (1) static world generation models, which construct 3D environments without active agents, and (2)…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Yitong Wang , Fangyun Wei , Hongyang Zhang , Bo Dai , Yan Lu

Grasping manipulation is a fundamental mode for human interaction with daily life objects. The synthesis of grasping motion is also greatly demanded in many applications such as animation and robotics. In objects grasping research field,…

机器人学 · 计算机科学 2024-10-04 Quanquan Shao , Yi Fang

We build the first system to address the problem of reconstructing in-scene object manipulation from a monocular RGB video. It is challenging due to ill-posed scene reconstruction, ambiguous hand-object depth, and the need for physically…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Dixuan Lin , Tianyou Wang , Zhuoyang Pan , Yufu Wang , Lingjie Liu , Kostas Daniilidis

We introduce Being-H0, a dexterous Vision-Language-Action model (VLA) trained on large-scale human videos. Existing VLAs struggle with complex manipulation tasks requiring high dexterity and generalize poorly to novel scenarios and tasks,…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Hao Luo , Yicheng Feng , Wanpeng Zhang , Sipeng Zheng , Ye Wang , Haoqi Yuan , Jiazheng Liu , Chaoyi Xu , Qin Jin , Zongqing Lu

Domain shifts, such as appearance changes, are a key challenge in real-world applications of activity recognition models, which range from assistive robotics and smart homes to driver observation in intelligent vehicles. For example, while…

计算机视觉与模式识别 · 计算机科学 2023-11-27 Zdravko Marinov , David Schneider , Alina Roitberg , Rainer Stiefelhagen