中文
相关论文

相关论文: MoCha:End-to-End Video Character Replacement witho…

200 篇论文

Deep generative models have been recently extended to synthesizing 3D digital humans. However, previous approaches treat clothed humans as a single chunk of geometry without considering the compositionality of clothing and accessories. As a…

计算机视觉与模式识别 · 计算机科学 2023-05-30 Taeksoo Kim , Shunsuke Saito , Hanbyul Joo

Real-time rendering of human head avatars is a cornerstone of many computer graphics applications, such as augmented reality, video games, and films, to name a few. Recent approaches address this challenge with computationally efficient…

计算机视觉与模式识别 · 计算机科学 2024-09-19 Kartik Teotia , Hyeongwoo Kim , Pablo Garrido , Marc Habermann , Mohamed Elgharib , Christian Theobalt

Photo-realistic and controllable 3D avatars are crucial for various applications such as virtual and mixed reality (VR/MR), telepresence, gaming, and film production. Traditional methods for avatar creation often involve time-consuming…

A great challenge in video-language (VidL) modeling lies in the disconnection between fixed video representations extracted from image/video understanding models and downstream VidL data. Recent studies try to mitigate this disconnection…

计算机视觉与模式识别 · 计算机科学 2022-04-19 Tsu-Jui Fu , Linjie Li , Zhe Gan , Kevin Lin , William Yang Wang , Lijuan Wang , Zicheng Liu

Many vision applications require identity consistency beyond strict biometric recognition, especially under non-frontal views or when facial cues are missing. However, conventional face recognition models enforce intra-identity invariance,…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Yingfeng Wang , Yuxuan Xiao , Shengcai Liao

To date, little attention has been given to multi-view 3D human mesh estimation, despite real-life applicability (e.g., motion capture, sport analysis) and robustness to single-view ambiguities. Existing solutions typically suffer from poor…

计算机视觉与模式识别 · 计算机科学 2022-12-13 Xuan Gong , Liangchen Song , Meng Zheng , Benjamin Planche , Terrence Chen , Junsong Yuan , David Doermann , Ziyan Wu

Existing neural human rendering methods struggle with a single image input due to the lack of information in invisible areas and the depth ambiguity of pixels in visible areas. In this regard, we propose Monocular Neural Human Renderer…

计算机视觉与模式识别 · 计算机科学 2022-10-04 Hongsuk Choi , Gyeongsik Moon , Matthieu Armando , Vincent Leroy , Kyoung Mu Lee , Gregory Rogez

Vision-Language-Action (VLA) models enable generalist robotic manipulation but suffer from high inference latency. This bottleneck stems from the massive number of visual tokens processed by large language backbones. Existing methods either…

机器人学 · 计算机科学 2026-03-12 Yuquan Li , Lianjie Ma , Han Ding , Lijun Zhu

Head avatars animated by visual signals have gained popularity, particularly in cross-driving synthesis where the driver differs from the animated character, a challenging but highly practical approach. The recently presented MegaPortraits…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Nikita Drobyshev , Antoni Bigata Casademunt , Konstantinos Vougioukas , Zoe Landgraf , Stavros Petridis , Maja Pantic

Existing methodologies for animating portrait images face significant challenges, particularly in handling non-frontal perspectives, rendering dynamic objects around the portrait, and generating immersive, realistic backgrounds. In this…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Jiahao Cui , Hui Li , Yun Zhan , Hanlin Shang , Kaihui Cheng , Yuqi Ma , Shan Mu , Hang Zhou , Jingdong Wang , Siyu Zhu

Creating realistic animations of human faces with computer graphic models is still a challenging task. It is often solved either with tedious manual work or motion capture based techniques that require specialised and costly hardware.…

计算机视觉与模式识别 · 计算机科学 2021-04-13 Wolfgang Paier , Anna Hilsmann , Peter Eisert

Talking head generation with arbitrary identities and speech audio remains a crucial problem in the realm of the virtual metaverse. Recently, diffusion models have become a popular generative technique in this field with their strong…

图形学 · 计算机科学 2025-08-11 Xinyang Li , Gen Li , Zhihui Lin , Yichen Qian , GongXin Yao , Weinan Jia , Aowen Wang , Weihua Chen , Fan Wang

Recently, self-supervised pre-training has advanced Vision Transformers on various tasks w.r.t. different data modalities, e.g., image and 3D point cloud data. In this paper, we explore this learning paradigm for 3D mesh data analysis based…

计算机视觉与模式识别 · 计算机科学 2022-07-22 Yaqian Liang , Shanshan Zhao , Baosheng Yu , Jing Zhang , Fazhi He

Recognizing human activities in videos is challenging due to the spatio-temporal complexity and context-dependence of human interactions. Prior studies often rely on single input modalities, such as RGB or skeletal data, limiting their…

计算机视觉与模式识别 · 计算机科学 2024-09-05 Tuyen Tran , Thao Minh Le , Hung Tran , Truyen Tran

The field of video generation has expanded significantly in recent years, with controllable and compositional video generation garnering considerable interest. Most methods rely on leveraging annotations such as text, objects' bounding…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Aram Davtyan , Sepehr Sameni , Björn Ommer , Paolo Favaro

Dynamic facial emotion is essential for believable AI-generated avatars, yet most systems remain visually static, limiting their use in simulations like virtual training for investigative interviews with abused children. We present a…

人机交互 · 计算机科学 2025-07-09 Pegah Salehi , Sajad Amouei Sheshkal , Vajira Thambawita , Michael A. Riegler , Pål Halvorsen

Multi-robot collaboration in large-scale environments with limited-sized teams and without external infrastructure is challenging, since the software framework required to support complex tasks must be robust to unreliable and intermittent…

机器人学 · 计算机科学 2023-09-29 Fernando Cladera , Zachary Ravichandran , Ian D. Miller , M. Ani Hsieh , C. J. Taylor , Vijay Kumar

We propose a self-supervised shared encoder model that achieves strong results on several visual, language and multimodal benchmarks while being data, memory and run-time efficient. We make three key contributions. First, in contrast to…

计算机视觉与模式识别 · 计算机科学 2023-04-13 Rakesh Chada , Zhaoheng Zheng , Pradeep Natarajan

Audio-driven talking head generation necessitates seamless integration of audio and visual data amidst the challenges posed by diverse input portraits and intricate correlations between audio and facial motions. In response, we propose a…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Ziqi Zhou , Weize Quan , Hailin Shi , Wei Li , Lili Wang , Dong-Ming Yan

Medical image re-identification (MedReID) is under-explored so far, despite its critical applications in personalized healthcare and privacy protection. In this paper, we introduce a thorough benchmark and a unified model for this problem.…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Yuan Tian , Kaiyuan Ji , Rongzhao Zhang , Yankai Jiang , Chunyi Li , Xiaosong Wang , Guangtao Zhai