English
Related papers

Related papers: Seedance 2.0: Advancing Video Generation for World…

200 papers

Diffusion probabilistic models (DPMs) and their extensions have emerged as competitive generative models yet confront challenges of efficient sampling. We propose a new bilateral denoising diffusion model (BDDM) that parameterizes both the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-28 Max W. Y. Lam , Jun Wang , Dan Su , Dong Yu

Significant development of communication technology over the past few years has motivated research in multi-modal summarization techniques. A majority of the previous works on multi-modal summarization focus on text and images. In this…

Information Retrieval · Computer Science 2020-05-20 Anubhav Jangra , Sriparna Saha , Adam Jatowt , Mohammad Hasanuzzaman

Generating consistent long videos is a complex challenge: while diffusion-based generative models generate visually impressive short clips, extending them to longer durations often leads to memory bottlenecks and long-term inconsistency. In…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Wenqi Ouyang , Zeqi Xiao , Danni Yang , Yifan Zhou , Shuai Yang , Lei Yang , Jianlou Si , Xingang Pan

Producing long, coherent video sequences with stable 3D structure remains a major challenge, particularly in streaming scenarios. Motivated by this, we introduce Endless World, a real-time framework for infinite, 3D-consistent video…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Ke Zhang , Yiqun Mei , Jiacong Xu , Vishal M. Patel

Streaming voice conversion has become increasingly popular for its potential in real-time applications. The recently proposed DualVC 2 has achieved robust and high-quality streaming voice conversion with a latency of about 180ms.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-13 Ziqian Ning , Shuai Wang , Pengcheng Zhu , Zhichao Wang , Jixun Yao , Lei Xie , Mengxiao Bi

Generative AI is reshaping the media landscape, enabling unprecedented capabilities in video creation, personalization, and scalability. This paper presents a comprehensive SWOT analysis of Metas Movie Gen, a cutting-edge generative AI…

Artificial Intelligence · Computer Science 2024-12-06 Abul Ehtesham , Saket Kumar , Aditi Singh , Tala Talaei Khoei

Recent advancements in 4D generation have demonstrated its remarkable capability in synthesizing photorealistic renderings of dynamic 3D scenes. However, despite achieving impressive visual performance, almost all existing methods overlook…

Sound · Computer Science 2026-03-02 Siyi Xie , Hanxin Zhu , Xinyi Chen , Tianyu He , Xin Li , Zhibo Chen

In Video on Demand (VoD) scenarios, traditional codecs are the industry standard due to their high decoding efficiency. However, they suffer from severe quality degradation under low bandwidth conditions. While emerging generative neural…

Image and Video Processing · Electrical Eng. & Systems 2026-02-20 Liming Liu , Jiangkai Wu , Haoyang Wang , Peiheng Wang , Zongming Guo , Xinggong Zhang

In recent years, diffusion models have made remarkable strides in text-to-video generation, sparking a quest for enhanced control over video outputs to more accurately reflect user intentions. Traditional efforts predominantly focus on…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Mingxiao Li , Bo Wan , Marie-Francine Moens , Tinne Tuytelaars

Large speech generation models are evolving from single-speaker, short sentence synthesis to multi-speaker, long conversation geneartion. Current long-form speech generation models are predominately constrained to dyadic, turn-based…

While traditional and neural video codecs (NVCs) have achieved remarkable rate-distortion performance, improving perceptual quality at low bitrates remains challenging. Some NVCs incorporate perceptual or adversarial objectives but still…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Naifu Xue , Zhaoyang Jia , Jiahao Li , Bin Li , Zihan Zheng , Yuan Zhang , Yan Lu

While text-to-video diffusion models have advanced significantly, creating coherent long-form content remains unreliable due to stochastic sampling artifacts. This necessitates generating multiple candidates, yet verifying them creates a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Daewon Yoon , Hyeongseok Lee , Wonsik Shin , Sangyu Han , Nojun Kwak

Human video generation is a dynamic and rapidly evolving task that aims to synthesize 2D human body video sequences with generative models given control conditions such as text, audio, and pose. With the potential for wide-ranging…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Wentao Lei , Jinting Wang , Fengji Ma , Guanjie Huang , Li Liu

Recent advancements have established Diffusion Transformers (DiTs) as a dominant framework in generative modeling. Building on this success, Lumina-Next achieves exceptional performance in the generation of photorealistic images with…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Dongyang Liu , Shicheng Li , Yutong Liu , Zhen Li , Kai Wang , Xinyue Li , Qi Qin , Yufei Liu , Yi Xin , Zhongyu Li , Bin Fu , Chenyang Si , Yuewen Cao , Conghui He , Ziwei Liu , Yu Qiao , Qibin Hou , Hongsheng Li , Peng Gao

Audio is indispensable for real-world video, yet generation models have largely overlooked audio components. Current approaches to producing audio-visual content often rely on cascaded pipelines, which increase cost, accumulate errors, and…

In recent years, with the realistic generation results and a wide range of personalized applications, diffusion-based generative models gain huge attention in both visual and audio generation areas. Compared to the considerable advancements…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Shiqi Yang , Zhi Zhong , Mengjie Zhao , Shusuke Takahashi , Masato Ishii , Takashi Shibuya , Yuki Mitsufuji

Video-to-music (V2M) generation aims to create music that aligns with visual content. However, two main challenges persist in existing methods: (1) the lack of explicit rhythm modeling hinders audiovisual temporal alignments; (2)…

Sound · Computer Science 2025-11-13 Shulei Ji , Zihao Wang , Jiaxing Yu , Xiangyuan Yang , Shuyu Li , Songruoyao Wu , Kejun Zhang

Diffusion-based image compression has demonstrated impressive perceptual performance. However, it suffers from two critical drawbacks: (1) excessive decoding latency due to multi-step sampling, and (2) poor fidelity resulting from…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Zheng Chen , Mingde Zhou , Jinpei Guo , Jiale Yuan , Yifei Ji , Yulun Zhang

Recent advancement in Generative Adversarial Networks in speech synthesis domain[3],[2] have shown, that it's possible to train GANs [8] in a reliable manner for high quality coherent waveform generation from mel-spectograms. We propose…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-16 Luka Chkhetiani , Levan Bejanidze