English
Related papers

Related papers: ActAvatar: Temporally-Aware Precise Action Control…

200 papers

Recent talking avatar generation models have made strides in achieving realistic and accurate lip synchronization with the audio, but often fall short in controlling and conveying detailed expressions and emotions of the avatar, making the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Yuchi Wang , Junliang Guo , Jianhong Bai , Runyi Yu , Tianyu He , Xu Tan , Xu Sun , Jiang Bian

Recently, 2D speaking avatars have increasingly participated in everyday scenarios due to the fast development of facial animation techniques. However, most existing works neglect the explicit control of human bodies. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Jiazhi Guan , Quanwei Yang , Kaisiyuan Wang , Hang Zhou , Shengyi He , Zhiliang Xu , Haocheng Feng , Errui Ding , Jingdong Wang , Hongtao Xie , Youjian Zhao , Ziwei Liu

Diffusion models have shown impressive potential on talking head generation. While plausible appearance and talking effect are achieved, these methods still suffer from temporal, 3D or expression inconsistency due to the error accumulation…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Haijie Yang , Zhenyu Zhang , Hao Tang , Jianjun Qian , Jian Yang

Significant progress has been made in audio-driven human animation, while most existing methods focus mainly on facial movements, limiting their ability to create full-body animations with natural synchronization and fluidity. They also…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Qijun Gan , Ruizi Yang , Jianke Zhu , Shaofei Xue , Steven Hoi

Avatar video generation models have achieved remarkable progress in recent years. However, prior work exhibits limited efficiency in generating long-duration high-resolution videos, suffering from temporal drifting, quality degradation, and…

Open-Vocabulary Temporal Action Detection (OV-TAD) aims to classify and localize action segments in untrimmed videos for unseen categories. Previous methods rely solely on global alignment between label-level semantics and visual features,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Sa Zhu , Wanqian Zhang , Lin Wang , Xiaohua Chen , Chenxu Cui , Jinchao Zhang , Bo Li

Temporal modeling is crucial for various video learning tasks. Most recent approaches employ either factorized (2D+1D) or joint (3D) spatial-temporal operations to extract temporal contexts from the input frames. While the former is more…

Computer Vision and Pattern Recognition · Computer Science 2023-01-03 Yizhou Zhao , Zhenyang Li , Xun Guo , Yan Lu

Existing video avatar models have demonstrated impressive capabilities in scenarios such as talking, public speaking, and singing. However, the majority of these methods exhibit limited alignment with respect to text instructions,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Ruikui Wang , Jinheng Feng , Lang Tian , Huaishao Luo , Chaochao Li , Liangbo Zhou , Huan Zhang , Youzheng Wu , Xiaodong He

Audio-visual automatic speech recognition (AV-ASR) is an extension of ASR that incorporates visual cues, often from the movements of a speaker's mouth. Unlike works that simply focus on the lip motion, we investigate the contribution of…

Computer Vision and Pattern Recognition · Computer Science 2022-06-16 Valentin Gabeur , Paul Hongsuck Seo , Arsha Nagrani , Chen Sun , Karteek Alahari , Cordelia Schmid

Real-time synthesis of high-fidelity 3D character motion from audio is a pivotal component for next-generation interactive avatars and virtual assistants. However, most existing approaches are limited to offline processing of complete audio…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Bohong Chen , Yumeng Li , Yinglin Xu , Youyi Zheng , Yanlin Weng , Kun Zhou

In human dialogue, nonverbal information such as nodding and facial expressions is as crucial as verbal information, and spoken dialogue systems are also expected to express such nonverbal behaviors. We focus on nodding, which is critical…

Human-Computer Interaction · Computer Science 2025-08-05 Kazushi Kato , Koji Inoue , Divesh Lala , Keiko Ochi , Tatsuya Kawahara

In Audio-Visual Navigation (AVN), agents must locate sound sources in unseen 3D environments using visual and auditory cues. However, existing methods often struggle with generalization in unseen scenarios, as they tend to overfit to…

Sound · Computer Science 2026-04-08 Jia Li , Yinfeng Yu

Recent advances in generative video models have enabled the creation of high-quality videos based on natural language prompts. However, these models frequently lack fine-grained temporal control, meaning they do not allow users to specify…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Shira Schiber , Ofir Lindenbaum , Idan Schwartz

Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain. Recent UDA methods based on Vision Transformers (ViTs) have achieved strong performance through attention-based…

Machine Learning · Computer Science 2025-06-24 Zelin Zang , Fei Wang , Liangyu Li , Jinlin Wu , Chunshui Zhao , Zhen Lei , Baigui Sun

This paper presents a novel spatiotemporal transformer network that introduces several original components to detect actions in untrimmed videos. First, the multi-feature selective semantic attention model calculates the correlations…

Computer Vision and Pattern Recognition · Computer Science 2024-05-15 Matthew Korban , Peter Youngs , Scott T. Acton

Creating a realistic animatable avatar from a single static portrait remains challenging. Existing approaches often struggle to capture subtle facial expressions, the associated global body movements, and the dynamic background. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Mengchao Wang , Qiang Wang , Fan Jiang , Yaqi Fan , Yunpeng Zhang , Yonggang Qi , Kun Zhao , Mu Xu

Recently, text-guided digital portrait editing has attracted more and more attentions. However, existing methods still struggle to maintain consistency across time, expression, and view or require specific data prerequisites. To solve these…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Haiyao Xiao , Chenglai Zhong , Xuan Gao , Yudong Guo , Juyong Zhang

Weakly-supervised audio-visual video parsing (AVVP) seeks to detect audible, visible, and audio-visual events without temporal annotations. Previous work has emphasized refining global predictions through contrastive or collaborative…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Yaru Chen , Ruohao Guo , Liting Gao , Yang Xiang , Qingyu Luo , Zhenbo Li , Wenwu Wang

Self-attention learns pairwise interactions to model long-range dependencies, yielding great improvements for video action recognition. In this paper, we seek a deeper understanding of self-attention for temporal modeling in videos. We…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Bo He , Xitong Yang , Zuxuan Wu , Hao Chen , Ser-Nam Lim , Abhinav Shrivastava

Audio-driven human animation technology is widely used in human-computer interaction, and the emergence of diffusion models has further advanced its development. Currently, most methods rely on multi-stage generation and intermediate…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 S. Z. Zhou , Y. B. Wang , J. F. Wu , T. Hu , J. N. Zhang
‹ Prev 1 2 3 10 Next ›