English
Related papers

Related papers: DynFOA: Generating First-Order Ambisonics with Con…

200 papers

We present Diff3F as a simple, robust, and class-agnostic feature descriptor that can be computed for untextured input shapes (meshes or point clouds). Our method distills diffusion features from image foundational models onto input shapes.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-04 Niladri Shekhar Dutt , Sanjeev Muralikrishnan , Niloy J. Mitra

Can machines recording an audio-visual scene produce realistic, matching audio-visual experiences at novel positions and novel view directions? We answer it by studying a new task -- real-world audio-visual scene synthesis -- and a…

Computer Vision and Pattern Recognition · Computer Science 2023-10-17 Susan Liang , Chao Huang , Yapeng Tian , Anurag Kumar , Chenliang Xu

Advancements in 3D scene reconstruction have transformed 2D images from the real world into 3D models, producing realistic 3D results from hundreds of input photos. Despite great success in dense-view reconstruction scenarios, rendering a…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Fangfu Liu , Wenqiang Sun , Hanyang Wang , Yikai Wang , Haowen Sun , Junliang Ye , Jun Zhang , Yueqi Duan

Recent joint audio-visual diffusion models achieve remarkable generation quality but suffer from high latency due to their bidirectional attention dependencies, hindering real-time applications. We propose OmniForcing, the first framework…

Multimedia · Computer Science 2026-03-16 Yaofeng Su , Yuming Li , Zeyue Xue , Jie Huang , Siming Fu , Haoran Li , Ying Li , Zezhong Qian , Haoyang Huang , Nan Duan

This study presents a novel method for generating music visualisers using diffusion models, combining audio input with user-selected artwork. The process involves two main stages: image generation and video creation. First, music captioning…

Multimedia · Computer Science 2024-12-10 Leonardo Pina , Yongmin Li

Driven by the emergence of Controllable Video Diffusion, existing Sim2Real methods for autonomous driving video generation typically rely on explicit intermediate representations to bridge the domain gap. However, these modalities face a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Xuyang Chen , Conglang Zhang , Chuanheng Fu , Zihao Yang , Kaixuan Zhou , Yizhi Zhang , Jianan He , Yanfeng Zhang , Mingwei Sun , Zengmao Wang , Zhen Dong , Xiaoxiao Long , Liqiu Meng

Large-scale diffusion generative models are greatly simplifying image, video and 3D asset creation from user-provided text prompts and images. However, the challenging problem of text-to-4D dynamic 3D scene generation with diffusion…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Yufeng Zheng , Xueting Li , Koki Nagano , Sifei Liu , Karsten Kreis , Otmar Hilliges , Shalini De Mello

Generating physically plausible human motion is crucial for applications such as character animation and virtual reality. Existing approaches often incorporate a simulator-based motion projection layer to the diffusion process to enforce…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Akihisa Watanabe , Jiawei Ren , Li Siyao , Yichen Peng , Erwin Wu , Edgar Simo-Serra

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency, and producing…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Longtao Zheng , Yifan Zhang , Hanzhong Guo , Jiachun Pan , Zhenxiong Tan , Jiahao Lu , Chuanxin Tang , Bo An , Shuicheng Yan

Synthesizing novel views from monocular videos of dynamic scenes remains a challenging problem. Scene-specific methods that optimize 4D representations with explicit motion priors often break down in highly dynamic regions where multi-view…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Thomas Tanay , Mohammed Brahimi , Michal Nazarczuk , Qingwen Zhang , Sibi Catley-Chandar , Arthur Moreau , Zhensong Zhang , Eduardo Pérez-Pellitero

The introduction of neural radiance fields has greatly improved the effectiveness of view synthesis for monocular videos. However, existing algorithms face difficulties when dealing with uncontrolled or lengthy scenarios, and require…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Kaichen Zhou , Jia-Xing Zhong , Sangyun Shin , Kai Lu , Yiyuan Yang , Andrew Markham , Niki Trigoni

A core capability for robot manipulation is reasoning over where and how to stably place objects in cluttered environments. Traditionally, robots have relied on object-specific, hand-crafted heuristics in order to perform such reasoning,…

Robotics · Computer Science 2023-10-27 Takuma Yoneda , Tianchong Jiang , Gregory Shakhnarovich , Matthew R. Walter

This paper introduces MIDI, a novel paradigm for compositional 3D scene generation from a single image. Unlike existing methods that rely on reconstruction or retrieval techniques or recent approaches that employ multi-stage…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Zehuan Huang , Yuan-Chen Guo , Xingqiao An , Yunhan Yang , Yangguang Li , Zi-Xin Zou , Ding Liang , Xihui Liu , Yan-Pei Cao , Lu Sheng

Real-world applications like video gaming and virtual reality often demand the ability to model 3D scenes that users can explore along custom camera trajectories. While significant progress has been made in generating 3D objects from text…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Tianyu Huang , Wangguandong Zheng , Tengfei Wang , Yuhao Liu , Zhenwei Wang , Junta Wu , Jie Jiang , Hui Li , Rynson W. H. Lau , Wangmeng Zuo , Chunchao Guo

With the rise of diffusion models, audio-video generation has been revolutionized. However, most existing methods rely on separate modules for each modality, with limited exploration of unified generative architectures. In addition, many…

Multimedia · Computer Science 2025-07-08 Lei Zhao , Linfeng Feng , Dongxu Ge , Rujin Chen , Fangqiu Yi , Chi Zhang , Xiao-Lei Zhang , Xuelong Li

Designing complex 3D scenes has been a tedious, manual process requiring domain expertise. Emerging text-to-3D generative models show great promise for making this task more intuitive, but existing approaches are limited to object-level…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Ryan Po , Gordon Wetzstein

In recent years there have been remarkable breakthroughs in image-to-video generation. However, the 3D consistency and camera controllability of generated frames have remained unsolved. Recent studies have attempted to incorporate camera…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Dejia Xu , Yifan Jiang , Chen Huang , Liangchen Song , Thorsten Gernoth , Liangliang Cao , Zhangyang Wang , Hao Tang

Diffusion Transformer, the backbone of Sora for video generation, successfully scales the capacity of diffusion models, pioneering new avenues for high-fidelity sequential data generation. Unlike static data such as images, sequential data…

Machine Learning · Computer Science 2025-02-05 Hengyu Fu , Zehao Dou , Jiawei Guo , Mengdi Wang , Minshuo Chen

Text-guided diffusion models have shown superior performance in image/video generation and editing. While few explorations have been performed in 3D scenarios. In this paper, we discuss three fundamental and interesting problems on this…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Gang Li , Heliang Zheng , Chaoyue Wang , Chang Li , Changwen Zheng , Dacheng Tao

Video composition is the core task of video editing. Although image composition based on diffusion models has been highly successful, it is not straightforward to extend the achievement to video object composition tasks, which not only…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Wei Wang , Yaosen Chen , Yuegen Liu , Qi Yuan , Shubin Yang , Yanru Zhang