English
Related papers

Related papers: Geo-Align: Video Generation Alignment via Metric G…

200 papers

Measuring alignment between language and vision is a fundamental challenge, especially as multimodal data becomes increasingly detailed and complex. Existing methods often rely on collecting human or AI preferences, which can be costly and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Hyojin Bahng , Caroline Chan , Fredo Durand , Phillip Isola

Video matting has traditionally been limited by the lack of high-quality ground-truth data. Most existing video matting datasets provide only human-annotated imperfect alpha and foreground annotations, which must be composited to background…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Yongtao Ge , Kangyang Xie , Guangkai Xu , Mingyu Liu , Li Ke , Longtao Huang , Hui Xue , Hao Chen , Chunhua Shen

Recent advances in diffusion-based and autoregressive video generation models have achieved remarkable visual realism. However, these models typically lack accurate physical alignment, failing to replicate real-world dynamics in object…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Tao Feng , Xianbing Zhao , Zhenhua Chen , Tien Tsin Wong , Hamid Rezatofighi , Gholamreza Haffari , Lizhen Qu

Recent progress in text-to-video generation has achieved remarkable realism, yet fine-grained control over camera motion and orientation remains elusive, especially with extreme trajectories (e.g., a 180-degree turnaround, or looking…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Frédéric Fortier-Chouinard , Yannick Hold-Geoffroy , Valentin Deschaintre , Matheus Gadelha , Jean-François Lalonde

This paper presents a novel approach for unsupervised video summarization using reinforcement learning (RL), addressing limitations like unstable adversarial training and reliance on heuristic-based reward functions. The method operates on…

Multimedia · Computer Science 2025-12-24 Mehryar Abbasi , Hadi Hadizadeh , Parvaneh Saeedi

Generating videos for visual storytelling can be a tedious and complex process that typically requires either live-action filming or graphics animation rendering. To bypass these challenges, our key idea is to utilize the abundance of…

Computer Vision and Pattern Recognition · Computer Science 2023-07-14 Yingqing He , Menghan Xia , Haoxin Chen , Xiaodong Cun , Yuan Gong , Jinbo Xing , Yong Zhang , Xintao Wang , Chao Weng , Ying Shan , Qifeng Chen

Recent text-to-video (T2V) diffusion models have made remarkable progress in generating high-quality videos. However, they often struggle to align with complex text prompts, particularly when multiple objects, attributes, or spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Daeun Lee , Jaehong Yoon , Jaemin Cho , Mohit Bansal

Recent advancements in video generation have substantially improved visual quality and temporal coherence, making these models increasingly appealing for applications such as autonomous driving, particularly in the context of driving…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Chun-Peng Chang , Chen-Yu Wang , Julian Schmidt , Holger Caesar , Alain Pagani

Generative models have made significant progress in synthesizing visual content, including images, videos, and 3D/4D structures. However, they are typically trained with surrogate objectives such as likelihood or reconstruction loss, which…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Yuanzhi Liang , Yijie Fang , Ke Hao , Rui Li , Ziqi Ni , Ruijie Su , Chi Zhang

Video generation has achieved remarkable progress with the introduction of diffusion models, which have significantly improved the quality of generated videos. However, recent research has primarily focused on scaling up model training,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Chenyang Si , Weichen Fan , Zhengyao Lv , Ziqi Huang , Yu Qiao , Ziwei Liu

Physical principles are fundamental to realistic visual simulation, but remain a significant oversight in transformer-based video generation. This gap highlights a critical limitation in rendering rigid body motion, a core tenet of…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Qiyuan Zhang , Biao Gong , Shuai Tan , Zheng Zhang , Yujun Shen , Xing Zhu , Yuyuan Li , Kelu Yao , Chunhua Shen , Changqing Zou

Motion generation, the task of synthesizing realistic motion sequences from various conditioning inputs, has become a central problem in computer vision, computer graphics, and robotics, with applications ranging from animation and virtual…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Aliasghar Khani , Arianna Rampini , Bruno Roy , Larasika Nadela , Noa Kaplan , Evan Atherton , Derek Cheung , Jacky Bibliowicz

Given a monocular video, the goal of video re-rendering is to generate views of the scene from a novel camera trajectory. Existing methods face two distinct challenges. Geometrically unconditioned models lack spatial awareness, leading to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Mingyang Xie , Numair Khan , Tianfu Wang , Naina Dhingra , Seonghyeon Nam , Haitao Yang , Zhuo Hui , Christopher Metzler , Andrea Vedaldi , Hamed Pirsiavash , Lei Luo

Anime video generation faces significant challenges due to the scarcity of anime data and unusual motion patterns, leading to issues such as motion distortion and flickering artifacts, which result in misalignment with human preferences.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Bingwen Zhu , Yudong Jiang , Baohan Xu , Siqian Yang , Mingyu Yin , Yidi Wu , Huyang Sun , Zuxuan Wu

Diffusion models and flow matching have demonstrated remarkable success in text-to-image generation. While many existing alignment methods primarily focus on fine-tuning pre-trained generative models to maximize a given reward function,…

Machine Learning · Statistics 2026-02-03 Yidong Ouyang , Liyan Xie , Hongyuan Zha , Guang Cheng

In this work, we rethink the approach to video super-resolution by introducing a method based on the Diffusion Posterior Sampling framework, combined with an unconditional video diffusion transformer operating in latent space. The video…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Zhihao Zhan , Wang Pang , Xiang Zhu , Yechao Bai

Estimating accurate and temporally consistent 3D human geometry from videos is a challenging problem in computer vision. Existing methods, primarily optimized for single images, often suffer from temporal inconsistencies and fail to capture…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Gwanghyun Kim , Xueting Li , Ye Yuan , Koki Nagano , Tianye Li , Jan Kautz , Se Young Chun , Umar Iqbal

Visual generative models have achieved remarkable progress in synthesizing photorealistic images and videos, yet aligning their outputs with human preferences across critical dimensions remains a persistent challenge. Though reinforcement…

World models simulate dynamic environments, enabling agents to interact with diverse input modalities. Although recent advances have improved the visual quality and temporal consistency of video world models, their ability of accurately…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Yang Ye , Tianyu He , Shuo Yang , Jiang Bian

Reference-to-video (R2V) generation is a controllable video synthesis paradigm that constrains the generation process using both text prompts and reference images, enabling applications such as personalized advertising and virtual try-on.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Lei Wang , YuXin Song , Ge Wu , Haocheng Feng , Hang Zhou , Jingdong Wang , Yaxing Wang , jian Yang