中文
相关论文

相关论文: SoulX-FlashHead: Oracle-guided Generation of Infin…

200 篇论文

Deploying massive diffusion models for real-time, infinite-duration, audio-driven avatar generation presents a significant engineering challenge, primarily due to the conflict between computational load and strict latency constraints.…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Le Shen , Qian Qiao , Tan Yu , Ke Zhou , Tianhang Yu , Yu Zhan , Zhenjie Wang , Ming Tao , Shunshun Yin , Siyuan Liu

Real-time text-driven joint audio-video avatar generation requires jointly synthesizing portrait video and speech with high fidelity and precise synchronization, yet existing audio-visual diffusion models remain too slow for interactive use…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Chunyu Li , Jiaye Li , Ruiqiao Mei , Haoyuan Xia , Hao Zhu , Jingdong Wang , Siyu Zhu

Diffusion-based models have gained wide adoption in the virtual human generation due to their outstanding expressiveness. However, their substantial computational requirements have constrained their deployment in real-time interactive…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Haojie Yu , Zhaonian Wang , Yihan Pan , Meng Cheng , Hao Yang , Chao Wang , Tao Xie , Xiaoming Xu , Xiaoming Wei , Xunliang Cai

Current motion-conditioned video generation methods suffer from prohibitive latency (minutes per video) and non-causal processing that prevents real-time interaction. We present MotionStream, enabling sub-second latency with up to 29 FPS…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Joonghyuk Shin , Zhengqi Li , Richard Zhang , Jun-Yan Zhu , Jaesik Park , Eli Shechtman , Xun Huang

Generating realistic, dyadic talking head video requires ultra-low latency. Existing chunk-based methods require full non-causal context windows, introducing significant delays. This high latency critically prevents the immediate,…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Bohong Chen , Haiyang Liu

Vision-language models (VLMs) could power real-time assistants and autonomous agents, but they face a critical challenge: understanding near-infinite video streams without escalating latency and memory usage. Processing entire videos with…

计算机视觉与模式识别 · 计算机科学 2025-10-13 Ruyi Xu , Guangxuan Xiao , Yukang Chen , Liuning He , Kelly Peng , Yao Lu , Song Han

We propose a multimodal-driven framework for high-fidelity long-term digital human animation termed $\textbf{Soul}$, which generates semantically coherent videos from a single-frame portrait image, text prompts, and audio, achieving precise…

Generative models are reshaping the live-streaming industry by redefining how content is created, styled, and delivered. Previous image-based streaming diffusion models have powered efficient and creative live streaming products but have…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Tianrui Feng , Zhi Li , Shuo Yang , Haocheng Xi , Muyang Li , Xiuyu Li , Lvmin Zhang , Keting Yang , Kelly Peng , Song Han , Maneesh Agrawala , Kurt Keutzer , Akio Kodaira , Chenfeng Xu

Diffusion models have recently advanced video restoration, but applying them to real-world video super-resolution (VSR) remains challenging due to high latency, prohibitive computation, and poor generalization to ultra-high resolutions. Our…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Junhao Zhuang , Shi Guo , Xin Cai , Xiaohui Li , Yihao Liu , Chun Yuan , Tianfan Xue

A versatile video depth estimation model should (1) be accurate and consistent across frames, (2) produce high-resolution depth maps, and (3) support real-time streaming. We propose FlashDepth, a method that satisfies all three…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Gene Chou , Wenqi Xian , Guandao Yang , Mohamed Abdelfattah , Bharath Hariharan , Noah Snavely , Ning Yu , Paul Debevec

Benefiting from the advancements in large language models and cross-modal alignment, existing multi-modal video understanding methods have achieved prominent performance in offline scenario. However, online video streams, as one of the most…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Haoji Zhang , Yiqin Wang , Yansong Tang , Yong Liu , Jiashi Feng , Jifeng Dai , Xiaojie Jin

Audio-driven talking head generation has drawn much attention in recent years, and many efforts have been made in lip-sync, expressive facial expressions, natural head pose generation, and high video quality. However, no model has yet led…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Xusen Sun , Longhao Zhang , Hao Zhu , Peng Zhang , Bang Zhang , Xinya Ji , Kangneng Zhou , Daiheng Gao , Liefeng Bo , Xun Cao

In this work, we introduce the first autoregressive framework for real-time, audio-driven portrait animation, a.k.a, talking head. Beyond the challenge of lengthy animation times, a critical challenge in realistic talking head generation…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Dingcheng Zhen , Shunshun Yin , Shiyang Qin , Hou Yi , Ziwei Zhang , Siyuan Liu , Gan Qi , Ming Tao

Audio-driven avatar interaction demands real-time, streaming, and infinite-length generation -- capabilities fundamentally at odds with the sequential denoising and long-horizon drift of current diffusion models. We present Live Avatar, an…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Yubo Huang , Hailong Guo , Fangtai Wu , Weiqiang Wang , Shifeng Zhang , Shijie Huang , Qijun Gan , Lin Liu , Sirui Zhao , Enhong Chen , Jiaming Liu , Steven Hoi

Recent advancements in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment. While recent methods…

音频与语音处理 · 电气工程与系统科学 2025-06-04 Huadai Liu , Jialei Wang , Rongjie Huang , Yang Liu , Heng Lu , Zhou Zhao , Wei Xue

Recent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Zhenjie Liu , Jianzhang Lu , Renjie Lu , Cong Liang , Shangfei Wang

Most of the existing video face super-resolution (VFSR) methods are trained and evaluated on VoxCeleb1, which is designed specifically for speaker identification and the frames in this dataset are of low quality. As a consequence, the VFSR…

图像与视频处理 · 电气工程与系统科学 2022-05-10 Liangbin Xie. Xintao Wang , Honglun Zhang , Chao Dong , Ying Shan

Recent advances in diffusion models have driven remarkable progress in image generation. However, the generation process remains computationally intensive, and users often need to iteratively refine prompts to achieve the desired results,…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Yi Wei , Shunpu Tang , Liang Zhao , Qiangian Yang

Current diffusion-based portrait animation models predominantly focus on enhancing visual quality and expression realism, while overlooking generation latency and real-time performance, which restricts their application range in the live…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Zhiyuan Li , Chi-Man Pun , Chen Fang , Jue Wang , Xiaodong Cun

Existing DiT-based audio-driven avatar generation methods have achieved considerable progress, yet their broader application is constrained by limitations such as high computational overhead and the inability to synthesize long-duration…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Chaochao Li , Ruikui Wang , Liangbo Zhou , Jinheng Feng , Huaishao Luo , Huan Zhang , Youzheng Wu , Xiaodong He
‹ 上一页 1 2 3 10 下一页 ›