中文
相关论文

相关论文: OHTA: One-shot Hand Avatar via Data-driven Implici…

200 篇论文

Bimanual robotic manipulation is a long-standing challenge of embodied intelligence due to its characteristics of dual-arm spatial-temporal coordination and high-dimensional action spaces. Previous studies rely on pre-defined action…

机器人学 · 计算机科学 2025-04-29 Huayi Zhou , Ruixiang Wang , Yunxin Tai , Yueci Deng , Guiliang Liu , Kui Jia

The ability to grasp objects, signal with gestures, and share emotion through touch all stem from the unique capabilities of human hands. Yet creating high-quality personalized hand avatars from images remains challenging due to complex…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Zicong Fan , Edoardo Remelli , David Dimond , Fadime Sener , Liuhao Ge , Bugra Tekin , Cem Keskin , Shreyas Hampali

We present FHAvatar, a novel framework for reconstructing 3D Gaussian avatars with composable face and hair components from an arbitrary number of views. Unlike previous approaches that couple facial and hair representations within a…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Yujie Sun , Zhuoqiang Cai , Chaoyue Niu , Jianchuan Chen , Zhiwen Chen , Chengfei Lv , Fan Wu

Fine-tuning advanced diffusion models for high-quality image stylization usually requires large training datasets and substantial computational resources, hindering their practical applicability. We propose Ada-Adapter, a novel framework…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Jia Liu , Changlin Li , Qirui Sun , Jiahui Ming , Chen Fang , Jue Wang , Bing Zeng , Shuaicheng Liu

Today's Mixed Reality head-mounted displays track the user's head pose in world space as well as the user's hands for interaction in both Augmented Reality and Virtual Reality scenarios. While this is adequate to support user input, it…

计算机视觉与模式识别 · 计算机科学 2022-07-29 Jiaxi Jiang , Paul Streli , Huajian Qiu , Andreas Fender , Larissa Laich , Patrick Snape , Christian Holz

Creating high-fidelity 3D head avatars has always been a research hotspot, but it remains a great challenge under lightweight sparse view setups. In this paper, we propose HHAvatar represented by controllable 3D Gaussians for high-fidelity…

计算机视觉与模式识别 · 计算机科学 2024-11-21 Zhanfeng Liao , Yuelang Xu , Zhe Li , Qijing Li , Boyao Zhou , Ruifeng Bai , Di Xu , Hongwen Zhang , Yebin Liu

We present LiftAvatar, a new paradigm that completes sparse monocular observations in kinematic space (e.g., facial expressions and head pose) and uses the completed signals to drive high-fidelity avatar animation. LiftAvatar is a…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Hualiang Wei , Shunran Jia , Jialun Liu , Wenhui Li

In this work, we advance the neural head avatar technology to the megapixel resolution while focusing on the particularly challenging task of cross-driving synthesis, i.e., when the appearance of the driving image is substantially different…

计算机视觉与模式识别 · 计算机科学 2023-03-29 Nikita Drobyshev , Jenya Chelishev , Taras Khakhulin , Aleksei Ivakhnenko , Victor Lempitsky , Egor Zakharov

Vision-language-action models (VLAs) trained on large-scale robotic datasets have demonstrated strong performance on manipulation tasks, including bimanual tasks. However, because most public datasets focus on single-arm demonstrations,…

机器人学 · 计算机科学 2026-02-24 Hokyun Im , Euijin Jeong , Andrey Kolobov , Jianlong Fu , Youngwoon Lee

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (GHOI) remains an open…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Youliang Zhang , Zhengguang Zhou , Zhentao Yu , Ziyao Huang , Teng Hu , Sen Liang , Guozhen Zhang , Ziqiao Peng , Shunkai Li , Yi Chen , Zixiang Zhou , Yuan Zhou , Qinglin Lu , Xiu Li

We propose a method for synthesizing edited photo-realistic digital avatars with text instructions. Given a short monocular RGB video and text instructions, our method uses an image-conditioned diffusion model to edit one head image and…

计算机视觉与模式识别 · 计算机科学 2023-06-06 Shaoxu Li

The appearance of a human in clothing is driven not only by the pose but also by its temporal context, i.e., motion. However, such context has been largely neglected by existing monocular human modeling methods whose neural networks often…

计算机视觉与模式识别 · 计算机科学 2023-12-29 Hansol Lee , Junuk Cha , Yunhoe Ku , Jae Shin Yoon , Seungryul Baek

To make 3D human avatars widely available, we must be able to generate a variety of 3D virtual humans with varied identities and shapes in arbitrary poses. This task is challenging due to the diversity of clothed body shapes, their complex…

计算机视觉与模式识别 · 计算机科学 2022-04-14 Xu Chen , Tianjian Jiang , Jie Song , Jinlong Yang , Michael J. Black , Andreas Geiger , Otmar Hilliges

Multi-view volumetric rendering techniques have recently shown great potential in modeling and synthesizing high-quality head avatars. A common approach to capture full head dynamic performances is to track the underlying geometry using a…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Kartik Teotia , Mallikarjun B R , Xingang Pan , Hyeongwoo Kim , Pablo Garrido , Mohamed Elgharib , Christian Theobalt

We introduce AvatarForge, a framework for generating animatable 3D human avatars from text or image inputs using AI-driven procedural generation. While diffusion-based methods have made strides in general 3D object generation, they struggle…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Xinhang Liu , Yu-Wing Tai , Chi-Keung Tang

We address the challenging task of anticipating human-object interaction in first person videos. Most existing methods ignore how the camera wearer interacts with the objects, or simply consider body motion as a separate modality. In…

计算机视觉与模式识别 · 计算机科学 2020-07-21 Miao Liu , Siyu Tang , Yin Li , James Rehg

Dexterity is often seen as a cornerstone of complex manipulation. Humans are able to perform a host of skills with their hands, from making food to operating tools. In this paper, we investigate these challenges, especially in the case of…

机器人学 · 计算机科学 2023-12-13 Aditya Kannan , Kenneth Shaw , Shikhar Bahl , Pragna Mannam , Deepak Pathak

Progress in human behavior modeling involves understanding both implicit, early-stage perceptual behavior, such as human attention, and explicit, later-stage behavior, such as subjective preferences or likes. Yet most prior research has…

Speech-driven 3D facial animation has been widely explored, with applications in gaming, character animation, virtual reality, and telepresence systems. State-of-the-art methods deform the face topology of the target actor to sync the input…

计算机视觉与模式识别 · 计算机科学 2023-01-03 Balamurugan Thambiraja , Ikhsanul Habibie , Sadegh Aliakbarian , Darren Cosker , Christian Theobalt , Justus Thies

Creating animatable 3D avatars from a single image remains challenging due to style limitations (realistic, cartoon, anime) and difficulties in handling accessories or hairstyles. While 3D diffusion models advance single-view reconstruction…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Tingting Liao , Yujian Zheng , Adilbek Karmanov , Liwen Hu , Leyang Jin , Yuliang Xiu , Hao Li