English
Related papers

Related papers: EchoTorrent: Towards Swift, Sustained, and Streami…

200 papers

Benefiting from the advances in large language models and cross-modal alignment, existing multimodal large language models have achieved prominent performance in image and short video understanding. However, the understanding of long videos…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Haoji Zhang , Yiqin Wang , Yansong Tang , Yong Liu , Jiashi Feng , Xiaojie Jin

Audio-driven talking-head generation has advanced rapidly with diffusion-based generative models, yet producing temporally coherent videos with fine-grained motion control remains challenging. We propose DEMO, a flow-matching generative…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Peiyin Chen , Zhuowei Yang , Hui Feng , Sheng Jiang , Rui Yan

Echocardiography (ECHO) is essential for cardiac assessments, but its video quality and interpretation heavily relies on manual expertise, leading to inconsistent results from clinical and portable devices. ECHO video generation offers a…

Image and Video Processing · Electrical Eng. & Systems 2024-10-15 Yiwei Li , Sekeun Kim , Zihao Wu , Hanqi Jiang , Yi Pan , Pengfei Jin , Sifan Song , Yucheng Shi , Tianming Liu , Quanzheng Li , Xiang Li

Diffusion-based talking head generation has achieved remarkable visual quality, yet scaling it to long-term videos remains challenging. The widely adopted chunk-wise paradigm introduces two fundamental failures: (1) temporal-spatial…

Machine Learning · Computer Science 2026-05-12 Yuxin Lu , Jiayang Sun , Guibo Zhu , Min Cao

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency, and producing…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Longtao Zheng , Yifan Zhang , Hanzhong Guo , Jiachun Pan , Zhenxiong Tan , Jiahao Lu , Chuanxin Tang , Bo An , Shuicheng Yan

State-of-the-art video generative models typically learn the distribution of video latents in the VAE space and map them to pixels using a VAE decoder. While this approach can generate high-quality videos, it suffers from slow convergence…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Jianhong Bai , Xiaoshi Wu , Xintao Wang , Xiao Fu , Yuanxing Zhang , Qinghe Wang , Xiaoyu Shi , Menghan Xia , Zuozhu Liu , Haoji Hu , Pengfei Wan , Kun Gai

Generative models have been widely applied to world modeling for environment simulation and future state prediction. With advancements in autonomous driving, there is a growing demand not only for high-fidelity video generation under…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Tianrui Zhang , Yichen Liu , Zilin Guo , Yuxin Guo , Jingcheng Ni , Chenjing Ding , Dan Xu , Lewei Lu , Zehuan Wu

Talking head generation with arbitrary identities and speech audio remains a crucial problem in the realm of the virtual metaverse. Recently, diffusion models have become a popular generative technique in this field with their strong…

Graphics · Computer Science 2025-08-11 Xinyang Li , Gen Li , Zhihui Lin , Yichen Qian , GongXin Yao , Weinan Jia , Aowen Wang , Weihua Chen , Fan Wang

Audio generation has attracted significant attention. Despite remarkable enhancement in audio quality, existing models overlook diversity evaluation. This is partially due to the lack of a systematic sound class diversity framework and a…

Sound · Computer Science 2024-07-19 Baihan Li , Zeyu Xie , Xuenan Xu , Yiwei Guo , Ming Yan , Ji Zhang , Kai Yu , Mengyue Wu

Diffusion models have demonstrated great success in text-to-video (T2V) generation. However, existing methods may face challenges when handling complex (long) video generation scenarios that involve multiple objects or dynamic changes in…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Ye Tian , Ling Yang , Haotian Yang , Yuan Gao , Yufan Deng , Jingmin Chen , Xintao Wang , Zhaochen Yu , Xin Tao , Pengfei Wan , Di Zhang , Bin Cui

Latent Video Diffusion Models (LVDMs) rely on Variational Autoencoders (VAEs) to compress videos into compact latent representations. For continuous Variational Autoencoders (VAEs), achieving higher compression rates is desirable; yet, the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Yubo Dong , Linchao Zhu

In natural face-to-face interaction, participants seamlessly alternate between speaking and listening, producing facial behaviors (FBs) that are finely informed by long-range context and naturally exhibit contextual appropriateness and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Xiangyu Kong , Xiaoyu Jin , Yihan Pan , Haoqin Sun , Hengde Zhu , Xiaoming Xu , Xiaoming Wei , Lu Liu , Siyang Song

Generative modeling has recently achieved remarkable success across text, image, and audio domains, demonstrating powerful capabilities for unified representation learning. However, audio generation models still face challenges in terms of…

Sound · Computer Science 2025-10-31 Chengwei Liu , Haoyin Yan , Shaofei Xue , Xiaotao Liang , Yinghao Liu , Zheng Xue , Gang Song , Boyang Zhou

Human-Centric Video Generation (HCVG) methods seek to synthesize human videos from multimodal inputs, including text, image, and audio. Existing methods struggle to effectively coordinate these heterogeneous modalities due to two…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Liyang Chen , Tianxiang Ma , Jiawei Liu , Bingchuan Li , Zhuowei Chen , Lijie Liu , Xu He , Gen Li , Qian He , Zhiyong Wu

Recent advances in text-to-video generation have achieved impressive performance on short clips, yet evaluating long-form generation under complex textual inputs remains a significant challenge. In response to this challenge, we present…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Xiangqing Zheng , Chengyue Wu , Kehai Chen , Min Zhang

Distilled video generation models offer fast and efficient synthesis but struggle with motion customization when guided by reference videos, especially under training-free settings. Existing training-free methods, originally designed for…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Jintao Rong , Xin Xie , Xinyi Yu , Linlin Ou , Xinyu Zhang , Chunhua Shen , Dong Gong

Compared to images, videos better reflect real-world acquisition and possess valuable temporal cues. However, existing multi-sensor fusion research predominantly integrates complementary context from multiple images rather than videos due…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Linfeng Tang , Yeda Wang , Meiqi Gong , Zizhuo Li , Yuxin Deng , Xunpeng Yi , Chunyu Li , Han Xu , Hao Zhang , Jiayi Ma

There has been growing sentiment recently that modern large multimodal models (LMMs) have addressed most of the key challenges related to short video comprehension. As a result, both academia and industry are gradually shifting their…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Jianrui Zhang , Mu Cai , Yong Jae Lee

Recent Video-to-Audio (V2A) generation relies on extracting semantic and temporal features from video to condition generative models. Training these models from scratch is resource intensive. Consequently, leveraging foundation models (FMs)…

Computer Vision and Pattern Recognition · Computer Science 2025-09-08 Gehui Chen , Guan'an Wang , Xiaowen Huang , Jitao Sang

Motion, speech, and sound effects are fundamental elements of human-centric videos, yet their heterogeneous temporal characteristics make joint generation highly challenging. Existing audio-video generation models often fail to maintain…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Shihao Cheng , Jiaxu Zhang , Quanyue Song , Shansong Liu , Zhizhi Guo , Xiaolei Zhang , Chi Zhang , Xuelong Li , Zhigang Tu