中文
相关论文

相关论文: SpriteHand: Real-Time Versatile Hand-Object Intera…

200 篇论文

We study the problem of making 3D scene reconstructions interactive by asking the following question: can we predict the sounds of human hands physically interacting with a scene? First, we record a video of a human manipulating objects…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Yiming Dou , Wonseok Oh , Yuqing Luo , Antonio Loquercio , Andrew Owens

Predicting the dynamics of interacting objects is essential for both humans and intelligent systems. However, existing approaches are limited to simplified, toy settings and lack generalizability to complex, real-world environments. Recent…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Rick Akkerman , Haiwen Feng , Michael J. Black , Dimitrios Tzionas , Victoria Fernández Abrevaya

Hand-object interaction understanding and the barely addressed novel view synthesis are highly desired in the immersive communication, whereas it is challenging due to the high deformation of hand and heavy occlusions between hand and…

计算机视觉与模式识别 · 计算机科学 2023-08-23 Wentian Qu , Zhaopeng Cui , Yinda Zhang , Chenyu Meng , Cuixia Ma , Xiaoming Deng , Hongan Wang

Most RGB-based hand-object reconstruction methods rely on object templates, while template-free methods typically assume full object visibility. This assumption often breaks in real-world settings, where fixed camera viewpoints and static…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Shibo Wang , Haonan He , Maria Parelli , Christoph Gebhardt , Zicong Fan , Jie Song

Humans frequently grasp, manipulate, and move objects. Interactive systems assist humans in these tasks, enabling applications in Embodied AI, human-robot interaction, and virtual reality. However, current methods in hand-object synthesis…

机器人学 · 计算机科学 2025-03-10 Sammy Christen

Synthesizing human motions in 3D environments, particularly those with complex activities such as locomotion, hand-reaching, and human-object interaction, presents substantial demands for user-defined waypoints and stage transitions. These…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Nan Jiang , Zimo He , Zi Wang , Hongjie Li , Yixin Chen , Siyuan Huang , Yixin Zhu

Real-time synthesis of physically plausible human interactions remains a critical challenge for immersive VR/AR systems and humanoid robotics. While existing methods demonstrate progress in kinematic motion generation, they often fail to…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Kaiyang Ji , Ye Shi , Zichen Jin , Kangyi Chen , Lan Xu , Yuexin Ma , Jingyi Yu , Jingya Wang

We present BimArt, a novel generative approach for synthesizing 3D bimanual hand interactions with articulated objects. Unlike prior works, we do not rely on a reference grasp, a coarse hand trajectory, or separate modes for grasping and…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Wanyue Zhang , Rishabh Dabral , Vladislav Golyanik , Vasileios Choutas , Eduardo Alvarado , Thabo Beeler , Marc Habermann , Christian Theobalt

With the advancement of interactive video generation, diffusion models have increasingly demonstrated their potential as world models. However, existing approaches still struggle to simultaneously achieve memory-enabled long-term temporal…

Achieving a balance between high-fidelity visual quality and low-latency streaming remains a formidable challenge in audio-driven portrait generation. Existing large-scale models often suffer from prohibitive computational costs, while…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Tan Yu , Qian Qiao , Le Shen , Ke Zhou , Jincheng Hu , Dian Sheng , Bo Hu , Haoming Qin , Jun Gao , Changhai Zhou , Shunshun Yin , Siyuan Liu

Extended reality (XR) demands generative models that respond to users' tracked real-world motion, yet current video world models accept only coarse control signals such as text or keyboard input, limiting their utility for embodied…

计算机视觉与模式识别 · 计算机科学 2026-02-23 Linxi Xie , Lisong C. Sun , Ashley Neall , Tong Wu , Shengqu Cai , Gordon Wetzstein

The rapid progress in artificial intelligence-generated content (AIGC), especially with diffusion models, has significantly advanced development of high-quality video generation. However, current video diffusion models exhibit demanding…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Zheng Zhan , Yushu Wu , Yifan Gong , Zichong Meng , Zhenglun Kong , Changdi Yang , Geng Yuan , Pu Zhao , Wei Niu , Yanzhi Wang

Speech-driven 3D facial animation aims to generate realistic lip movements and facial expressions for 3D head models from arbitrary audio clips. Although existing diffusion-based methods are capable of producing natural motions, their slow…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Xuangeng Chu , Nabarun Goswami , Ziteng Cui , Hanqin Wang , Tatsuya Harada

Gesture recognition research, unlike NLP, continues to face acute data scarcity, with progress constrained by the need for costly human recordings or image processing approaches that cannot generate authentic variability in the gestures…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Hassan Ali , Doreen Jirak , Luca Müller , Stefan Wermter

Achieving streaming, fine-grained control over the outputs of autoregressive video diffusion models remains challenging, making it difficult to ensure that they consistently align with user expectations. To bridge this gap, we propose…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Junbao Zhou , Yuan Zhou , Kesen Zhao , Qingshan Xu , Beier Zhu , Richang Hong , Hanwang Zhang

Vision-based human-to-robot handover is an important and challenging task in human-robot interaction. Recent work has attempted to train robot policies by interacting with dynamic virtual humans in simulated environments, where the policies…

机器人学 · 计算机科学 2025-01-03 Sammy Christen , Lan Feng , Wei Yang , Yu-Wei Chao , Otmar Hilliges , Jie Song

Understanding and synthesizing realistic 3D hand-object interactions (HOI) is critical for applications ranging from immersive AR/VR to dexterous robotics. Existing methods struggle with generalization, performing well on closed-set objects…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Zhenhao Zhang , Ye Shi , Lingxiao Yang , Suting Ni , Qi Ye , Jingya Wang

Existing large-scale video generation models are computationally intensive, preventing adoption in real-time and interactive applications. In this work, we propose autoregressive adversarial post-training (AAPT) to transform a pre-trained…

计算机视觉与模式识别 · 计算机科学 2025-10-03 Shanchuan Lin , Ceyuan Yang , Hao He , Jianwen Jiang , Yuxi Ren , Xin Xia , Yang Zhao , Xuefeng Xiao , Lu Jiang

Modern video generative models based on diffusion models can produce very realistic clips, but they are computationally inefficient, often requiring minutes of GPU time for just a few seconds of video. This inefficiency poses a critical…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Jieying Chen , Jeffrey Hu , Joan Lasenby , Ayush Tewari

Panoramic video generation aims to synthesize 360-degree immersive videos, holding significant importance in the fields of VR, world models, and spatial intelligence. Existing works fail to synthesize high-quality panoramic videos due to…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Zixun Fang , Kai Zhu , Zhiheng Liu , Yu Liu , Wei Zhai , Yang Cao , Zheng-Jun Zha