English
Related papers

Related papers: Auto-CARD: Efficient and Robust Codec Avatar Drivi…

200 papers

We introduce a novel representation for efficient classical rendering of photorealistic 3D face avatars. Leveraging recent advances in radiance fields anchored to parametric face models, our approach achieves controllable volumetric…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Safa C. Medin , Gengyan Li , Ziqian Bai , Ruofei Du , Leonhard Helminger , Yinda Zhang , Stephan J. Garbin , Philip L. Davidson , Gregory W. Wornell , Thabo Beeler , Abhimitra Meka

What is a good visual representation for autonomous agents? We address this question in the context of semantic visual navigation, which is the problem of a robot finding its way through a complex environment to a target object, e.g. go to…

Computer Vision and Pattern Recognition · Computer Science 2019-07-04 Arsalan Mousavian , Alexander Toshev , Marek Fiser , Jana Kosecka , Ayzaan Wahid , James Davidson

Quantitative metrics are central to evaluating computer vision (CV) models, but they often fail to capture real-world performance due to protocol inconsistencies and ground-truth noise. While visual perception studies can complement these…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Ashkan Ganj , Yiqin Zhao , Tian Guo

Audio-driven facial animation presents an effective solution for animating digital avatars. In this paper, we detail the technical aspects of NVIDIA Audio2Face-3D, including data acquisition, network architecture, retargeting methodology,…

We propose the notion of Attention-Aware Visualizations (AAVs) that track the user's perception of a visual representation over time and feed this information back to the visualization. Such context awareness is particularly useful for…

Human-Computer Interaction · Computer Science 2025-01-16 Arvind Srinivasan , Johannes Ellemose , Peter W. S. Butcher , Panagiotis D. Ritsos , Niklas Elmqvist

Modern video codecs including the newly developed AOM/AV1 utilize hybrid coding techniques to remove spatial and temporal redundancy. However, efficient exploitation of statistical dependencies measured by a mean squared error (MSE) does…

Image and Video Processing · Electrical Eng. & Systems 2018-04-26 Di Chen , Chichen Fu , Fengqing Zhu

Autonomous vehicles and Advanced Driving Assistance Systems (ADAS) have the potential to radically change the way we travel. Many such vehicles currently rely on segmentation and object detection algorithms to detect and track objects…

Computer Vision and Pattern Recognition · Computer Science 2023-07-28 Ravi Kakaiya , Rakshith Sathish , Ramanathan Sethuraman , Debdoot Sheet

Despite rapid advances in autonomous driving technology, current autonomous vehicles (AVs) lack effective bidirectional human-machine communication, limiting their ability to personalize the riding experience and recover from uncertain or…

Robotics · Computer Science 2025-11-17 Zhipeng Bao , Qianwen Li

Autonomous vehicles (AVs) must be both safe and trustworthy to gain social acceptance and become a viable option for everyday public transportation. Explanations about the system behaviour can increase safety and trust in AVs.…

Logic in Computer Science · Computer Science 2025-11-19 Dominik Grundt , Ishan Saxena , Malte Petersen , Bernd Westphal , Eike Möhlmann

Face reenactment methods attempt to restore and re-animate portrait videos as realistically as possible. Existing methods face a dilemma in quality versus controllability: 2D GAN-based methods achieve higher image quality but suffer in…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Lizhen Wang , Xiaochen Zhao , Jingxiang Sun , Yuxiang Zhang , Hongwen Zhang , Tao Yu , Yebin Liu

Egocentric motion capture with a head-mounted body-facing stereo camera is crucial for VR and AR applications but presents significant challenges such as heavy occlusions and limited annotated real-world data. Existing methods rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Andrea Boscolo Camiletto , Jian Wang , Eduardo Alvarado , Rishabh Dabral , Thabo Beeler , Marc Habermann , Christian Theobalt

Reconstructing personalized animatable head avatars has significant implications in the fields of AR/VR. Existing methods for achieving explicit face control of 3D Morphable Models (3DMM) typically rely on multi-view images or videos of a…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Haoyu Ma , Tong Zhang , Shanlin Sun , Xiangyi Yan , Kun Han , Xiaohui Xie

While Neural Processing Units (NPUs) offer high theoretical efficiency for edge AI, state-of-the-art Vision--Language Models (VLMs) tailored for GPUs often falter on these substrates. We attribute this hardware-model mismatch to two primary…

Computation and Language · Computer Science 2025-12-09 Wei Chen , Liangmin Wu , Yunhai Hu , Zhiyuan Li , Zhiyuan Cheng , Yicheng Qian , Lingyue Zhu , Zhipeng Hu , Luoyi Liang , Qiang Tang , Zhen Liu , Han Yang

Deep generative models, and particularly facial animation schemes, can be used in video conferencing applications to efficiently compress a video through a sparse set of keypoints, without the need to transmit dense motion vectors. While…

Multimedia · Computer Science 2022-07-28 Goluck Konuko , Stéphane Lathuilière , Giuseppe Valenzise

With advancements in computer vision and deep learning, video-based human action recognition (HAR) has become practical. However, due to the complexity of the computation pipeline, running HAR on live video streams incurs excessive delays…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Ruiqi Wang , Zichen Wang , Peiqi Gao , Mingzhen Li , Jaehwan Jeong , Yihang Xu , Yejin Lee , Carolyn M. Baum , Lisa Tabor Connor , Chenyang Lu

Neural implicit fields are powerful for representing 3D scenes and generating high-quality novel views, but it remains challenging to use such implicit representations for creating a 3D human avatar with a specific identity and artistic…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Ruixiang Jiang , Can Wang , Jingbo Zhang , Menglei Chai , Mingming He , Dongdong Chen , Jing Liao

With the rapid advancement of autonomous driving, deploying Vision-Language Models (VLMs) to enhance perception and decision-making has become increasingly common. However, the real-time application of VLMs is hindered by high latency and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-13 Lianming Huang , Haibo Hu , Yufei Cui , Jiacheng Zuo , Shangyu Wu , Nan Guan , Chun Jason Xue

Autonomous vehicles (AVs) promise efficient, clean and cost-effective transportation systems, but their reliance on sensors, wireless communications, and decision-making systems makes them vulnerable to cyberattacks and physical threats.…

Cryptography and Security · Computer Science 2026-04-15 Chieh Tsai , Murad Mehrab Abrar , Salim Hariri

Human pose estimation has achieved significant progress in recent years. However, most of the recent methods focus on improving accuracy using complicated models and ignoring real-time efficiency. To achieve a better trade-off between…

Computer Vision and Pattern Recognition · Computer Science 2021-05-24 Lumin Xu , Yingda Guan , Sheng Jin , Wentao Liu , Chen Qian , Ping Luo , Wanli Ouyang , Xiaogang Wang

Active speaker detection (ASD) and virtual cinematography (VC) can significantly improve the remote user experience of a video conference by automatically panning, tilting and zooming of a video conferencing camera: users subjectively rate…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-26 Ross Cutler , Ramin Mehran , Sam Johnson , Cha Zhang , Adam Kirk , Oliver Whyte , Adarsh Kowdle