English
Related papers

Related papers: MOVA: Towards Scalable and Synchronized Video-Audi…

200 papers

Speech enhancement plays an essential role in various applications, and the integration of visual information has been demonstrated to bring substantial advantages. However, the majority of current research concentrates on the examination…

Sound · Computer Science 2025-04-03 Xinyuan Qian , Jiaran Gao , Yaodan Zhang , Qiquan Zhang , Hexin Liu , Leibny Paola Garcia , Haizhou Li

Video and audio content creation serves as the core technique for the movie industry and professional users. Recently, existing diffusion-based methods tackle video and audio generation separately, which hinders the technique transfer from…

Computer Vision and Pattern Recognition · Computer Science 2024-02-29 Yazhou Xing , Yingqing He , Zeyue Tian , Xintao Wang , Qifeng Chen

We introduce Motion-I2V, a novel framework for consistent and controllable image-to-video generation (I2V). In contrast to previous methods that directly learn the complicated image-to-video mapping, Motion-I2V factorizes I2V into two…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Xiaoyu Shi , Zhaoyang Huang , Fu-Yun Wang , Weikang Bian , Dasong Li , Yi Zhang , Manyuan Zhang , Ka Chun Cheung , Simon See , Hongwei Qin , Jifeng Dai , Hongsheng Li

The evolution of video generation from text, from animating MNIST to simulating the world with Sora, has progressed at a breakneck speed. Here, we systematically discuss how far text-to-video generation technology supports essential…

For immersive applications, the generation of binaural sound that matches its visual counterpart is crucial to bring meaningful experiences to people in a virtual environment. Recent studies have shown the possibility of using neural…

Sound · Computer Science 2023-05-22 Francesc Lluís , Vasileios Chatziioannou , Alex Hofmann

We introduce ImmerseDiffusion, an end-to-end generative audio model that produces 3D immersive soundscapes conditioned on the spatial, temporal, and environmental conditions of sound objects. ImmerseDiffusion is trained to generate…

Sound · Computer Science 2025-02-11 Mojtaba Heydari , Mehrez Souden , Bruno Conejo , Joshua Atkins

The imitation of voice, targeted on specific speech attributes such as timbre and speaking style, is crucial in speech generation. However, existing methods rely heavily on annotated data, and struggle with effectively disentangling timbre…

Sora is a text-to-video generative AI model, released by OpenAI in February 2024. The model is trained to generate videos of realistic or imaginative scenes from text instructions and show potential in simulating the physical world. Based…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Yixin Liu , Kai Zhang , Yuan Li , Zhiling Yan , Chujie Gao , Ruoxi Chen , Zhengqing Yuan , Yue Huang , Hanchi Sun , Jianfeng Gao , Lifang He , Lichao Sun

Audio-visual generation is rapidly advancing from short clips to minute-long content, while existing evaluation protocols remain largely confined to short-form settings. Existing benchmarks primarily focus on 5--10 second text-conditioned…

The content of visual and audio scenes is multi-faceted such that a video can be paired with various audio and vice-versa. Thereby, in video-to-audio generation task, it is imperative to introduce steering approaches for controlling the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Xiulong Liu , Kun Su , Eli Shlizerman

We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or independent modules for each modality, VINO uses a shared diffusion…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Junyi Chen , Tong He , Zhoujie Fu , Pengfei Wan , Kun Gai , Weicai Ye

Video generative models are increasingly used as world models for robotics, where a model generates a future visual rollout conditioned on the current observation and task instruction, and an inverse dynamics model (IDM) converts the…

Robotics · Computer Science 2026-03-25 Ruixiang Wang , Qingming Liu , Yueci Deng , Guiliang Liu , Zhen Liu , Kui Jia

Currently, high-quality, synchronized audio is synthesized from video and optional text inputs using various multi-modal joint learning frameworks. However, the precise alignment between the visual and generated audio domains remains far…

Sound · Computer Science 2025-03-31 Yunming Liang , Zihao Chen , Chaofan Ding , Xinhan Di

Numerous studies in the field of music generation have demonstrated impressive performance, yet virtually no models are able to directly generate music to match accompanying videos. In this work, we develop a generative music AI framework,…

Sound · Computer Science 2024-06-03 Jaeyong Kang , Soujanya Poria , Dorien Herremans

Autoregressive sequence modeling stands as the cornerstone of modern Generative AI, powering results across diverse modalities ranging from text generation to image generation. However, a fundamental limitation of this paradigm is the rigid…

Machine Learning · Computer Science 2026-02-02 Yangyan Li

State-of-the-art Text-to-Video (T2V) diffusion models can generate visually impressive results, yet they still frequently fail to compose complex scenes or follow logical temporal instructions. In this paper, we argue that many errors,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Mariam Hassan , Bastien Van Delft , Wuyang Li , Alexandre Alahi

In this paper, we propose a novel audio-driven talking head method capable of simultaneously generating highly expressive facial expressions and hand gestures. Unlike existing methods that focus on generating full-body or half-body poses,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Linrui Tian , Siqi Hu , Qi Wang , Bang Zhang , Liefeng Bo

Recent progress in unified models for image understanding and generation has been impressive, yet most approaches remain limited to single-modal generation conditioned on multiple modalities. In this paper, we present Mogao, a unified…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Chao Liao , Liyang Liu , Xun Wang , Zhengxiong Luo , Xinyu Zhang , Wenliang Zhao , Jie Wu , Liang Li , Zhi Tian , Weilin Huang

Conventional GAN-based models for talking head generation often suffer from limited quality and unstable training. Recent approaches based on diffusion models aimed to address these limitations and improve fidelity. However, they still face…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Seyeon Kim , Siyoon Jin , Jihye Park , Kihong Kim , Jiyoung Kim , Jisu Nam , Seungryong Kim

Recently, non-orthogonal multiple access (NOMA) has been proposed to achieve higher spectral efficiency over conventional orthogonal multiple access. Although it has the potential to meet increasing demands of video services, it is still…

Information Theory · Computer Science 2018-01-17 Xiaoda Jiang , Hancheng Lu , Chang Wen Chen