中文
相关论文

相关论文: CamViG: Camera Aware Image-to-Video Generation wit…

200 篇论文

The rapid evolution of video generation has enabled models to simulate complex physical dynamics and long-horizon causalities, positioning them as potential world simulators. However, a critical gap still remains between the theoretical…

图像与视频处理 · 电气工程与系统科学 2026-05-06 Muyang He , Hanzhong Guo , Junxiong Lin , Yizhou Yu

Generating realistic robotic manipulation videos is an important step toward unifying perception, planning, and action in embodied agents. While existing video diffusion models require large domain-specific datasets and struggle to…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Ye Pang

While image manipulation achieves tremendous breakthroughs (e.g., generating realistic faces) in recent years, video generation is much less explored and harder to control, which limits its applications in the real world. For instance,…

计算机视觉与模式识别 · 计算机科学 2019-08-08 Tsun-Hsuan Wang , Yen-Chi Cheng , Chieh Hubert Lin , Hwann-Tzong Chen , Min Sun

Recent methods have made notable progress in the visual quality of hand-object interaction video synthesis. However, most approaches rely on 2D control signals that lack spatial expressiveness and limit the utilization of synthetic 3D…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Mingjin Chen , Junhao Chen , Zhaoxin Fan , Yujian Lee , Zichen Dang , Lili Wang , Yawen Cui , Lap-Pui Chau , Yi Wang

We introduce Vid-CamEdit, a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view…

We propose PostCam, a framework for novel-view video generation that enables post-capture editing of camera trajectories in dynamic scenes. We find that existing video recapture methods suffer from suboptimal camera motion injection…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Yipeng Chen , Zhichao Ye , Zhenzhou Fang , Xinyu Chen , Xiaoyu Zhang , Jialing Liu , Nan Wang , Haomin Liu , Guofeng Zhang

This paper proposes the novel task of video generation conditioned on a SINGLE semantic label map, which provides a good balance between flexibility and quality in the generation process. Different from typical end-to-end approaches, which…

计算机视觉与模式识别 · 计算机科学 2019-03-12 Junting Pan , Chengyu Wang , Xu Jia , Jing Shao , Lu Sheng , Junjie Yan , Xiaogang Wang

Generating temporally coherent high fidelity video is an important milestone in generative modeling research. We make progress towards this milestone by proposing a diffusion model for video generation that shows very promising initial…

计算机视觉与模式识别 · 计算机科学 2022-06-24 Jonathan Ho , Tim Salimans , Alexey Gritsenko , William Chan , Mohammad Norouzi , David J. Fleet

While diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new challenges posed by T2MVid generation lie in the lack of…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Bing Li , Cheng Zheng , Wenxuan Zhu , Jinjie Mai , Biao Zhang , Peter Wonka , Bernard Ghanem

Diffusion based video generation has received extensive attention and achieved considerable success within both the academic and industrial communities. However, current efforts are mainly concentrated on single-objective or single-task…

计算机视觉与模式识别 · 计算机科学 2024-01-18 Ludan Ruan , Lei Tian , Chuanwei Huang , Xu Zhang , Xinyan Xiao

In this demo, we present VirtualConductor, a system that can generate conducting video from any given music and a single user's image. First, a large-scale conductor motion dataset is collected and constructed. Then, we propose Audio Motion…

计算机视觉与模式识别 · 计算机科学 2021-08-11 Delong Chen , Fan Liu , Zewen Li , Feng Xu

The burgeoning growth of video-to-music generation can be attributed to the ascendancy of multimodal generative models. However, there is a lack of literature that comprehensively combs through the work in this field. To fill this gap, this…

音频与语音处理 · 电气工程与系统科学 2025-12-23 Shulei Ji , Songruoyao Wu , Zihao Wang , Shuyu Li , Kejun Zhang

Despite having been studied to a great extent, the task of conditional generation of sequences of frames, or videos, remains extremely challenging. It is a common belief that a key step towards solving this task resides in modelling…

计算机视觉与模式识别 · 计算机科学 2021-09-09 David Kanaa , Vikram Voleti , Samira Ebrahimi Kahou , Christopher Pal

Existing music-driven 3D dance generation methods mainly concentrate on high-quality dance generation, but lack sufficient control during the generation process. To address these issues, we propose a unified framework capable of generating…

声音 · 计算机科学 2024-03-21 Ronghui Li , Yuqin Dai , Yachao Zhang , Jun Li , Jian Yang , Jie Guo , Xiu Li

Generating 3D scenes is still a challenging task due to the lack of readily available scene data. Most existing methods only produce partial scenes and provide limited navigational freedom. We introduce a practical and scalable solution…

图形学 · 计算机科学 2025-09-26 Zhaoyang Zhang , Yannick Hold-Geoffroy , Miloš Hašan , Ziwen Chen , Fujun Luan , Julie Dorsey , Yiwei Hu

We present 3DScenePrompt, a framework that generates the next video chunk from arbitrary-length input while enabling precise camera control and preserving scene consistency. Unlike methods conditioned on a single image or a short clip, we…

计算机视觉与模式识别 · 计算机科学 2025-12-16 JoungBin Lee , Jaewoo Jung , Jisang Han , Takuya Narihira , Kazumi Fukuda , Junyoung Seo , Sunghwan Hong , Yuki Mitsufuji , Seungryong Kim

User-generated cinematic creations are gaining popularity as our daily entertainment, yet it is a challenge to master cinematography for producing immersive contents. Many existing automatic methods focus on roughly controlling predefined…

多媒体 · 计算机科学 2024-05-24 Xinyi Wu , Haohong Wang , Aggelos K. Katsaggelos

Generative AI (GenAI) has significantly advanced the ease and flexibility of image creation. However, it remains a challenge to precisely control spatial compositions, including object arrangement and scene conditions. To bridge this gap,…

人机交互 · 计算机科学 2025-08-12 Runlin Duan , Yuzhao Chen , Rahul Jain , Yichen Hu , Jingyu Shi , Karthik Ramani

Given a monocular video, the goal of video re-rendering is to generate views of the scene from a novel camera trajectory. Existing methods face two distinct challenges. Geometrically unconditioned models lack spatial awareness, leading to…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Mingyang Xie , Numair Khan , Tianfu Wang , Naina Dhingra , Seonghyeon Nam , Haitao Yang , Zhuo Hui , Christopher Metzler , Andrea Vedaldi , Hamed Pirsiavash , Lei Luo

We are living in a flourishing era of digital media, where everyone has the potential to become a personal filmmaker. Current research on video generation suggests a promising avenue for controllable film creation in pixel space using…

计算机视觉与模式识别 · 计算机科学 2025-07-03 Weiliang Chen , Fangfu Liu , Diankun Wu , Haowen Sun , Jiwen Lu , Yueqi Duan