English
Related papers

Related papers: JoyStreamer: Unlocking Highly Expressive Avatars v…

200 papers

Recent advances in diffusion-based video generation have enabled photo-realistic short clips, but current methods still struggle to achieve multi-modal consistency when jointly generating whole-body motion and natural speech. Current…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Xinhan Di , Kristin Qi , Pengqian Yu

Text-to-Avatar generation has recently made significant strides due to advancements in diffusion models. However, most existing work remains constrained by limited diversity, producing avatars with subtle differences in appearance for a…

Computer Vision and Pattern Recognition · Computer Science 2024-02-28 Weijing Tao , Biwen Lei , Kunhao Liu , Shijian Lu , Miaomiao Cui , Xuansong Xie , Chunyan Miao

The rising demand for creating lifelike avatars in the digital realm has led to an increased need for generating high-quality human videos guided by textual descriptions and poses. We propose Dancing Avatar, designed to fabricate human…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Bosheng Qin , Wentao Ye , Qifan Yu , Siliang Tang , Yueting Zhuang

Audio-visual automatic speech recognition (AV-ASR) is an extension of ASR that incorporates visual cues, often from the movements of a speaker's mouth. Unlike works that simply focus on the lip motion, we investigate the contribution of…

Computer Vision and Pattern Recognition · Computer Science 2022-06-16 Valentin Gabeur , Paul Hongsuck Seo , Arsha Nagrani , Chen Sun , Karteek Alahari , Cordelia Schmid

Recently, animating portrait images using audio input is a popular task. Creating lifelike talking head videos requires flexible and natural movements, including facial and head dynamics, camera motion, realistic light and shadow effects.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Wenzhang Sun , Xiang Li , Donglin Di , Zhuding Liang , Qiyuan Zhang , Hao Li , Wei Chen , Jianxun Cui

We present SentiAvatar, a framework for building expressive interactive 3D digital humans, and use it to create SuSu, a virtual character that speaks, gestures, and emotes in real time. Achieving such a system remains challenging, as it…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Chuhao Jin , Rui Zhang , Qingzhe Gao , Haoyu Shi , Dayu Wu , Yichen Jiang , Yihan Wu , Ruihua Song

Audio-driven portrait animation has made significant advances with diffusion-based models, improving video quality and lipsync accuracy. However, the increasing complexity of these models has led to inefficiencies in training and inference,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Xuyang Cao , Guoxin Wang , Sheng Shi , Jun Zhao , Yang Yao , Jintao Fei , Minyu Gao , Pei Xie

We propose TextToon, a method to generate a drivable toonified avatar. Given a short monocular video sequence and a written instruction about the avatar style, our model can generate a high-fidelity toonified avatar that can be driven in…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Luchuan Song , Lele Chen , Celong Liu , Pinxin Liu , Chenliang Xu

Generating images with accurately represented text, especially in non-Latin languages, poses a significant challenge for diffusion models. Existing approaches, such as the integration of hint condition diagrams via auxiliary networks (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Chao Li , Chen Jiang , Xiaolong Liu , Jun Zhao , Guoxin Wang

In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not only need automatic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Zheng Qin , Ruobing Zheng , Yabing Wang , Tianqi Li , Zixin Zhu , Sanping Zhou , Ming Yang , Le Wang

Due to the increasing demand in films and games, synthesizing 3D avatar animation has attracted much attention recently. In this work, we present a production-ready text/speech-driven full-body animation synthesis system. Given the text and…

Graphics · Computer Science 2022-06-01 Wenlin Zhuang , Jinwei Qi , Peng Zhang , Bang Zhang , Ping Tan

In this paper, we present TalkingMachines -- an efficient framework that transforms pretrained video generation models into real-time, audio-driven character animators. TalkingMachines enables natural conversational experiences by…

Sound · Computer Science 2025-06-04 Chetwin Low , Weimin Wang

We present READ Avatars, a 3D-based approach for generating 2D avatars that are driven by audio input with direct and granular control over the emotion. Previous methods are unable to achieve realistic animation due to the many-to-many…

Computer Vision and Pattern Recognition · Computer Science 2023-03-02 Jack Saunders , Vinay Namboodiri

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (GHOI) remains an open…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Youliang Zhang , Zhengguang Zhou , Zhentao Yu , Ziyao Huang , Teng Hu , Sen Liang , Guozhen Zhang , Ziqiao Peng , Shunkai Li , Yi Chen , Zixiang Zhou , Yuan Zhou , Qinglin Lu , Xiu Li

Audio-driven human video generation has achieved remarkable success in monologue scenarios, largely driven by advancements in powerful video generation foundation models. Moving beyond monologues, authentic human communication is inherently…

Artificial Intelligence · Computer Science 2026-04-14 Yuzhe Weng , Haotian Wang , Xinyi Yu , Xiaoyan Wu , Haoran Xu , Shan He , Jun Du

Recent years have witnessed significant progress in audio-driven human animation. However, critical challenges remain in (i) generating highly dynamic videos while preserving character consistency, (ii) achieving precise emotion alignment…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Yi Chen , Sen Liang , Zixiang Zhou , Ziyao Huang , Yifeng Ma , Junshu Tang , Qin Lin , Yuan Zhou , Qinglin Lu

Recent advances in video diffusion models have significantly improved character animation techniques. However, current approaches rely on basic structural conditions such as DWPose or SMPL-X to animate character images, limiting their…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Muyao Niu , Mingdeng Cao , Yifan Zhan , Qingtian Zhu , Mingze Ma , Jiancheng Zhao , Yanhong Zeng , Zhihang Zhong , Xiao Sun , Yinqiang Zheng

While diffusion model for audio-driven avatar video generation have achieved notable process in synthesizing long sequences with natural audio-visual synchronization and identity consistency, the generation of music-performance videos with…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Jiahui Chen , Weida Wang , Runhua Shi , Huan Yang , Chaofan Ding , Zihao Chen

Since the beginning of the COVID-19 pandemic, remote conferencing and school-teaching have become important tools. The previous applications aim to save the commuting cost with real-time interactions. However, our application is going to…

Artificial Intelligence · Computer Science 2022-10-14 Aolan Sun , Xulong Zhang , Tiandong Ling , Jianzong Wang , Ning Cheng , Jing Xiao

Current talking avatars mostly generate co-speech gestures based on audio and text of the utterance, without considering the non-speaking motion of the speaker. Furthermore, previous works on co-speech gesture generation have designed…

Multimedia · Computer Science 2024-01-09 Sicheng Yang , Zunnan Xu , Haiwei Xue , Yongkang Cheng , Shaoli Huang , Mingming Gong , Zhiyong Wu