中文
相关论文

相关论文: FlashSign: Pose-Free Guidance for Efficient Sign L…

200 篇论文

This paper introduces EasyAnimate, an efficient and high quality video generation framework that leverages diffusion transformers to achieve high-quality video production, encompassing data processing, model training, and end-to-end…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Jiaqi Xu , Kunzhe Huang , Xinyi Zou , Yunkuo Chen , Bo Liu , MengLi Cheng , Jun Huang , Xing Shi

Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. While various sparse attention…

计算与语言 · 计算机科学 2026-03-09 Qihang Fan , Huaibo Huang , Zhiying Wu , Juqiu Wang , Bingning Wang , Ran He

Recent breakthroughs in text-to-image diffusion models have significantly advanced the generation of high-fidelity, photo-realistic images from textual descriptions. Yet, these models often struggle with interpreting spatial arrangements…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Jiaqi Liu , Tao Huang , Chang Xu

Current diffusion-based acceleration methods for long-portrait animation struggle to ensure identity (ID) consistency. This paper presents FlashPortrait, an end-to-end video diffusion transformer capable of synthesizing ID-preserving,…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Shuyuan Tu , Yueming Pan , Yinming Huang , Xintong Han , Zhen Xing , Qi Dai , Kai Qiu , Chong Luo , Zuxuan Wu

Image composition involves seamlessly integrating given objects into a specific visual context. Current training-free methods rely on composing attention weights from several samplers to guide the generator. However, since these weights are…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Yibin Wang , Weizhong Zhang , Jianwei Zheng , Cheng Jin

Sign language is a fundamental means of communication for the deaf and hard-of-hearing (DHH) community, enabling nuanced expression through gestures, facial expressions, and body movements. Despite its critical role in facilitating…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Alexander Brettmann , Jakob Grävinghoff , Marlene Rüschoff , Marie Westhues

Existing large-scale sign language resources typically provide supervision only at the level of raw video-text alignment and are often produced in laboratory settings. While such resources are important for semantic understanding, they do…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Sen Fang , Hongbin Zhong , Yanxin Zhang , Dimitris N. Metaxas

Recent advancements in text-to-image diffusion models have demonstrated remarkable success, yet they often struggle to fully capture the user's intent. Existing approaches using textual inputs combined with bounding boxes or region masks…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Seonho Lee , Jiho Choi , Seohyun Lim , Jiwook Kim , Hyunjung Shim

In recent years, video conferencing applications have become increasingly prevalent, relying heavily on high-speed internet connectivity. When such connectivity is lacking, users often default to audio-only communication, a mode that…

多媒体 · 计算机科学 2025-11-12 Panneer Selvam Santhalingam , Swann Thantsin , Ahmad Kamari , Parth Pathak , Kenneth DeHaan

The visual dubbing task aims to generate mouth movements synchronized with the driving audio, which has seen significant progress in recent years. However, two critical deficiencies hinder their wide application: (1) Audio-only driving…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Liyang Chen , Tianze Zhou , Xu He , Boshi Tang , Zhiyong Wu , Yang Huang , Yang Wu , Zhongqian Sun , Wei Yang , Helen Meng

Talking head generation is to generate video based on a given source identity and target motion. However, current methods face several challenges that limit the quality and controllability of the generated videos. First, the generated face…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Yue Gao , Yuan Zhou , Jinglu Wang , Xiao Li , Xiang Ming , Yan Lu

With the rapid advancement of diffusion-based generative models, portrait image animation has achieved remarkable results. However, it still faces challenges in temporally consistent video generation and fast sampling due to its iterative…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Taekyung Ki , Dongchan Min , Gyeongsu Chae

Recently, large-scale diffusion models, e.g., Stable diffusion and DallE2, have shown remarkable results on image synthesis. On the other hand, large-scale cross-modal pre-trained models (e.g., CLIP, ALIGN, and FILIP) are competent for…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Runhui Huang , Jianhua Han , Guansong Lu , Xiaodan Liang , Yihan Zeng , Wei Zhang , Hang Xu

Image diffusion models are trained on independently sampled static images. While this is the bedrock task protocol in generative modeling, capturing the temporal world through the lens of static snapshots is information-deficient by design.…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Juhun Lee , Simon S. Woo

Along with the explosion of large language models, improvements in speech synthesis, advancements in hardware, and the evolution of computer graphics, the current bottleneck in creating digital humans lies in generating character movements…

人机交互 · 计算机科学 2026-01-30 Thanh Hoang-Minh

Visual diffusion models achieve remarkable progress, yet they are typically trained at limited resolutions due to the lack of high-resolution data and constrained computation resources, hampering their ability to generate high-fidelity…

计算机视觉与模式识别 · 计算机科学 2025-08-22 Haonan Qiu , Ning Yu , Ziqi Huang , Paul Debevec , Ziwei Liu

We introduce the hfut-lmc team's solution to the SLRTP Sign Production Challenge. The challenge aims to generate semantically aligned sign language pose sequences from text inputs. To this end, we propose a Text-driven Diffusion Model (TDM)…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Jiayi He , Xu Wang , Ruobei Zhang , Shengeng Tang , Yaxiong Wang , Lechao Cheng

Controllable text-to-image (T2I) diffusion models have shown impressive performance in generating high-quality visual content through the incorporation of various conditions. Current methods, however, exhibit limited performance when guided…

计算机视觉与模式识别 · 计算机科学 2024-11-06 Jiajun Wang , Morteza Ghahremani , Yitong Li , Björn Ommer , Christian Wachinger

Text-based talking-head video editing aims to efficiently insert, delete, and substitute segments of talking videos through a user-friendly text editing approach. It is challenging because of \textbf{1)} generalizable talking-face…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Bo Han , Heqing Zou , Haoyang Li , Guangcong Wang , Chng Eng Siong

Scene text recognition (STR) suffers from challenges of either less realistic synthetic training data or the difficulty of collecting sufficient high-quality real-world data, limiting the effectiveness of trained models. Meanwhile, despite…

计算机视觉与模式识别 · 计算机科学 2025-09-11 Xingsong Ye , Yongkun Du , Yunbo Tao , Zhineng Chen