English
Related papers

Related papers: WISA: World Simulator Assistant for Physics-Aware …

200 papers

Despite their wide-spread success, Text-to-Image models (T2I) still struggle to produce images that are both aesthetically pleasing and faithful to the user's input text. We introduce DreamSync, a model-agnostic training algorithm by design…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Jiao Sun , Deqing Fu , Yushi Hu , Su Wang , Royi Rassin , Da-Cheng Juan , Dana Alon , Charles Herrmann , Sjoerd van Steenkiste , Ranjay Krishna , Cyrus Rashtchian

Generative models have demonstrated remarkable capability in synthesizing high-quality text, images, and videos. For video generation, contemporary text-to-video models exhibit impressive capabilities, crafting visually stunning videos.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Jay Zhangjie Wu , Guian Fang , Haoning Wu , Xintao Wang , Yixiao Ge , Xiaodong Cun , David Junhao Zhang , Jia-Wei Liu , Yuchao Gu , Rui Zhao , Weisi Lin , Wynne Hsu , Ying Shan , Mike Zheng Shou

The Visual Physics Analysis (VISPA) project integrates different aspects of physics analyses into a graphical development environment. It addresses the typical development cycle of (re-)designing, executing and verifying an analysis. The…

Data Analysis, Statistics and Probability · Physics 2012-09-10 H. -P. Bretz , M. Brodski , M. Erdmann , R. Fischer , A. Hinzmann , T. Klimkovich , D. Klingebiel , M. Komm , J. Lingemann , G. Müller , T. Münzer , M. Rieger , J. Steggemann , T. Winchen

One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack representations that effectively embed spatio-temporal properties of visual scenes at the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Nikita Araslanov , Martin Sundermeyer , Hidenobu Matsuki , David Joseph Tan , Federico Tombari

Diagnosis in histopathology requires a global whole slide images (WSIs) analysis, requiring pathologists to compound evidence from different WSI patches. The gigapixel scale of WSIs poses a challenge for histopathology multi-modal models.…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Mehmet Saygin Seyfioglu , Wisdom O. Ikezogwo , Fatemeh Ghezloo , Ranjay Krishna , Linda Shapiro

Inferring the physical properties of 3D scenes from visual information is a critical yet challenging task for creating interactive and realistic virtual worlds. While humans intuitively grasp material characteristics such as elasticity or…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Long Le , Ryan Lucas , Chen Wang , Chuhao Chen , Dinesh Jayaraman , Eric Eaton , Lingjie Liu

Advances in neural fields are enabling high-fidelity capture of the shape and appearance of dynamic 3D scenes. However, their capabilities lag behind those offered by conventional representations such as 2D videos because of algorithmic…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Cheng-You Lu , Peisen Zhou , Angela Xing , Chandradeep Pokhariya , Arnab Dey , Ishaan Shah , Rugved Mavidipalli , Dylan Hu , Andrew Comport , Kefan Chen , Srinath Sridhar

Text-to-video (T2V) synthesis has advanced rapidly, yet current evaluation metrics primarily capture visual quality and temporal consistency, offering limited insight into how synthetic videos perform in downstream tasks such as…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Zecheng Zhao , Selena Song , Tong Chen , Zhi Chen , Shazia Sadiq , Yadan Luo

With the rapid advancement of video generation models such as Sora, video quality assessment (VQA) is becoming increasingly crucial for selecting high-quality videos from large-scale datasets used in pre-training. Traditional VQA methods,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Yanyun Pu , Kehan Li , Zeyi Huang , Zhijie Zhong , Kaixiang Yang

In recent years, data-driven techniques have greatly advanced autonomous driving systems, but the need for rare and diverse training data remains a challenge, requiring significant investment in equipment and labor. World models, which…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Haiguang Wang , Daqi Liu , Hongwei Xie , Haisong Liu , Enhui Ma , Kaicheng Yu , Limin Wang , Bing Wang

The rapid growth of text-to-video (T2V) diffusion models has raised concerns about privacy, copyright, and safety due to their potential misuse in generating harmful or misleading content. These models are often trained on numerous…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Naen Xu , Jinghuai Zhang , Changjiang Li , Zhi Chen , Chunyi Zhou , Qingming Li , Tianyu Du , Shouling Ji

Vehicle-to-Everything (V2X) collaborative perception is crucial for autonomous driving. However, achieving high-precision V2X perception requires a significant amount of annotated real-world data, which can always be expensive and hard to…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Xianghao Kong , Wentao Jiang , Jinrang Jia , Yifeng Shi , Runsheng Xu , Si Liu

Generative world models are increasingly used for video generation, where learned simulators are expected to capture the physical rules that govern real-world dynamics. However, evaluating whether generated videos actually follow these…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Juyi Lin , Arash Akbari , Yumei He , Lin Zhao , Haichao Zhang , Arman Akbari , Xingchen Xu , Zoe Y. Lu , Enfu Nan , Hokin Deng , Edmund Yeh , Sarah Ostadabbas , Yun Fu , Jennifer Dy , Pu Zhao , Yanzhi Wang

Distilling interpretable physical laws from videos has led to expanded interest in the computer vision community recently thanks to the advances in deep learning, but still remains a great challenge. This paper introduces an end-to-end…

Computer Vision and Pattern Recognition · Computer Science 2022-05-04 Lele Luan , Yang Liu , Hao Sun

Modeling sounds emitted from physical object interactions is critical for immersive perceptual experiences in real and virtual worlds. Traditional methods of impact sound synthesis use physics simulation to obtain a set of physics…

Computer Vision and Pattern Recognition · Computer Science 2023-07-11 Kun Su , Kaizhi Qian , Eli Shlizerman , Antonio Torralba , Chuang Gan

State-of-the-art video generative models produce promising visual content yet often violate basic physics principles, limiting their utility. While some attribute this deficiency to insufficient physics understanding from pre-training, we…

This paper presents a versatile image-to-image visual assistant, PixWizard, designed for image generation, manipulation, and translation based on free-from language instructions. To this end, we tackle a variety of vision tasks into a…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Weifeng Lin , Xinyu Wei , Renrui Zhang , Le Zhuo , Shitian Zhao , Siyuan Huang , Huan Teng , Junlin Xie , Yu Qiao , Peng Gao , Hongsheng Li

Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods attempt to inject 3D priors via architectural modifications, they often incur high…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Weijie Wang , Xiaoxuan He , Youping Gu , Yifan Yang , Zeyu Zhang , Yefei He , Yanbo Ding , Xirui Hu , Donny Y. Chen , Zhiyuan He , Yuqing Yang , Bohan Zhuang

Achieving semantic alignment across diverse video generation conditions remains a significant challenge. Methods that rely on explicit structural guidance often enforce rigid spatial constraints that limit semantic flexibility, whereas…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Zexi Wu , Baolu Li , Jing Dai , Yiming Zhang , Yue Ma , Qinghe Wang , Xu Jia , Hongming Xu

Recovering analytical solutions of physical fields from visual observations is a fundamental yet underexplored capability for AI-assisted scientific reasoning. We study visual-to-symbolic analytical solution inference (ViSA) for…