English
Related papers

Related papers: PhysAlign: Physics-Coherent Image-to-Video Generat…

200 papers

Current video generation models produce high-quality aesthetic videos but often struggle to learn representations of real-world physics dynamics, resulting in artifacts such as unnatural object collisions, inconsistent gravity, and temporal…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Siddarth Nilol Kundur Satish , Devesh Jaiswal , Hongyu Chen , Abhishek Bakshi

Recent open-world 3D representation learning methods using Vision-Language Models (VLMs) to align 3D point cloud with image-text information have shown superior 3D zero-shot performance. However, CAD-rendered images for this alignment often…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Ye Mao , Junpeng Jing , Krystian Mikolajczyk

Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Hila Chefer , Uriel Singer , Amit Zohar , Yuval Kirstain , Adam Polyak , Yaniv Taigman , Lior Wolf , Shelly Sheynin

Despite the advances in text-to-image synthesis, particularly with diffusion models, generating visual instructions that require consistent representation and smooth state transitions of objects across sequential steps remains a formidable…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Quynh Phung , Songwei Ge , Jia-Bin Huang

We present xGen-VideoSyn-1, a text-to-video (T2V) generation model capable of producing realistic scenes from textual descriptions. Building on recent advancements, such as OpenAI's Sora, we explore the latent diffusion model (LDM)…

Image-to-video (I2V) generation aims to use the initial frame (alongside a text prompt) to create a video sequence. A grand challenge in I2V generation is to maintain visual consistency throughout the video: existing methods often struggle…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Weiming Ren , Huan Yang , Ge Zhang , Cong Wei , Xinrun Du , Wenhao Huang , Wenhu Chen

The rapid advancement of Embodied AI has led to an increasing demand for large-scale, high-quality real-world data. However, collecting such embodied data remains costly and inefficient. As a result, simulation environments have become a…

Recovering 3D scenes from sparse views is a challenging task due to its inherent ill-posed problem. Conventional methods have developed specialized solutions (e.g., geometry regularization or feed-forward deterministic model) to mitigate…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Hanyang Wang , Fangfu Liu , Jiawei Chi , Yueqi Duan

Discovering physical laws directly from high-dimensional visual data is a long-standing human pursuit but remains a formidable challenge for machines, representing a fundamental goal of scientific intelligence. This task is inherently…

Computational Engineering, Finance, and Science · Computer Science 2026-02-24 Ruikun Li , Jun Yao , Yingfan Hua , Shixiang Tang , Biqing Qi , Bin Liu , Wanli Ouyang , Yan Lu

Creating high-dynamic videos such as motion-rich actions and sophisticated visual effects poses a significant challenge in the field of artificial intelligence. Unfortunately, current state-of-the-art video generation methods, primarily…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Yan Zeng , Guoqiang Wei , Jiani Zheng , Jiaxin Zou , Yang Wei , Yuchen Zhang , Hang Li

Generating geometrically consistent videos remains an open challenge: text-to-video diffusion models trained on web-scale data treat geometry only implicitly, leading to object deformation, texture drift, and non-rigid backgrounds under…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Jan Ackermann , Shengqu Cai , Boyang Deng , Zhengfei Kuang , Songyou Peng , Gordon Wetzstein

While vision-language models (VLMs) have demonstrated promising capabilities in reasoning and planning for embodied agents, their ability to comprehend physical phenomena, particularly within structured 3D environments, remains severely…

Diffusion-based \textit{image-to-video} (I2V) generation has become a central direction in generative models by turning a reference image, with optional conditions, into a temporally coherent video. Compared with broader video generation…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Xianlong Wang , Wenbo Pan , Shijia Zhou , Ke Li , Yuqi Wang , Zeyu Ye , Hangtao Zhang , Leo Yu Zhang , Xiaohua Jia

Reconstructing and simulating dynamic 3D scenes with both visual realism and physical consistency remains a fundamental challenge. Existing neural representations, such as NeRFs and 3DGS, excel in appearance reconstruction but struggle to…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Dan Wang , Xinrui Cui , Serge Belongie , Ravi Ramamoorthi

Recent advancements in image-to-video (I2V) generation have shown promising performance in conventional scenarios. However, these methods still encounter significant challenges when dealing with complex scenes that require a deep…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Peng Liu , Xiaoming Ren , Fengkai Liu , Qingsong Xie , Quanlong Zheng , Yanhao Zhang , Haonan Lu , Yujiu Yang

Generating high-quality videos that synthesize desired realistic content is a challenging task due to their intricate high-dimensionality and complexity of videos. Several recent diffusion-based methods have shown comparable performance by…

Computer Vision and Pattern Recognition · Computer Science 2024-04-05 Kihong Kim , Haneol Lee , Jihye Park , Seyeon Kim , Kwanghee Lee , Seungryong Kim , Jaejun Yoo

Vision-Language Models (VLMs) are increasingly applied to robotic perception and manipulation, yet their ability to infer physical properties required for manipulation remains limited. In particular, estimating the mass of real-world…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Hisayuki Yokomizo , Taiki Miyanishi , Yan Gang , Shuhei Kurita , Nakamasa Inoue , Yusuke Iwasawa

Text-to-video (T2V) generation is a rapidly growing research area that aims to translate the scenes, objects, and actions within complex video text into a sequence of coherent visual frames. We present FlowZero, a novel framework that…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Yu Lu , Linchao Zhu , Hehe Fan , Yi Yang

We introduce Motion-I2V, a novel framework for consistent and controllable image-to-video generation (I2V). In contrast to previous methods that directly learn the complicated image-to-video mapping, Motion-I2V factorizes I2V into two…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Xiaoyu Shi , Zhaoyang Huang , Fu-Yun Wang , Weikang Bian , Dasong Li , Yi Zhang , Manyuan Zhang , Ka Chun Cheung , Simon See , Hongwei Qin , Jifeng Dai , Hongsheng Li

Single-image-to-3D models typically follow a sequential generation and reconstruction workflow. However, intermediate multi-view images synthesized by pre-trained generation models often lack cross-view consistency (CVC), significantly…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Xinyue Liang , Zhiyuan Ma , Lingchen Sun , Yanjun Guo , Lei Zhang