English
Related papers

Related papers: Pixel2Phys: Distilling Governing Laws from Visual …

200 papers

Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering. Meanwhile, a smaller body of work addresses…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Ali Athar , Xueqing Deng , Liang-Chieh Chen

By providing substantial amounts of data and standardized evaluation protocols, datasets in computer vision have helped fuel advances across all areas of visual recognition. But even in light of breakthrough results on recent benchmarks, it…

Computer Vision and Pattern Recognition · Computer Science 2018-07-06 Brandon RichardWebster , Samuel E. Anthony , Walter J. Scheirer

Big-data-based artificial intelligence (AI) supports profound evolution in almost all of science and technology. However, modeling and forecasting multi-physical systems remain a challenge due to unavoidable data scarcity and noise.…

Machine Learning · Computer Science 2022-02-08 Pengpeng Shi , Zhi Zeng , Tianshou Liang

Rigid body interactions are fundamental to numerous scientific disciplines, but remain challenging to simulate due to their abrupt nonlinear nature and sensitivity to complex, often unknown environmental factors. These challenges call for…

Machine Learning · Computer Science 2025-07-28 Amaury Wei , Olga Fink

Partial Differential Equations (PDEs) have long been recognized as powerful tools for image processing and analysis, providing a framework to model and exploit structural and geometric properties inherent in visual data. Over the years,…

Image and Video Processing · Electrical Eng. & Systems 2024-12-17 Alejandro Garnung Menéndez

There is a growing interest in developing strong biomedical vision-language models. A popular approach to achieve robust representations is to use web-scale scientific data. However, current biomedical vision-language pretraining typically…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Kun Yuan , Min Woo Sun , Zhen Chen , Alejandro Lozano , Xiangteng He , Shi Li , Nassir Navab , Xiaoxiao Sun , Nicolas Padoy , Serena Yeung-Levy

Visual understanding goes well beyond object recognition. With one glance at an image, we can effortlessly imagine the world beyond the pixels: for instance, we can infer people's actions, goals, and mental states. While this task is easy…

Computer Vision and Pattern Recognition · Computer Science 2019-03-27 Rowan Zellers , Yonatan Bisk , Ali Farhadi , Yejin Choi

Discovering explicit physical laws has traditionally depended on human intuition and domain expertise. Recent advances in artificial intelligence, particularly large language models (LLMs), offer a new route to accelerate this process by…

Materials Science · Physics 2026-01-30 Bo Hu , Siyu Liu , Beilin Ye , Yun Hao , Yanhui Liu , Yang Lu , Ju Li , David J. Srolovitz , Tongqi Wen

Symmetries play a central role in physics, organizing dynamics, constraining interactions, and determining the effective number of physical degrees of freedom. In parallel, modern artificial intelligence methods have demonstrated a…

High Energy Physics - Phenomenology · Physics 2026-02-03 Veronica Sanz

Large Language Models (LLMs) have achieved remarkable progress on advanced reasoning tasks such as mathematics and coding competitions. Meanwhile, physics, despite being both reasoning-intensive and essential to real-world understanding,…

Computation and Language · Computer Science 2025-10-20 Shenghe Zheng , Qianjia Cheng , Junchi Yao , Mengsong Wu , Haonan He , Ning Ding , Yu Cheng , Shuyue Hu , Lei Bai , Dongzhan Zhou , Ganqu Cui , Peng Ye

Recent breakthroughs in Vision-Language (V&L) joint research have achieved remarkable results in various text-driven tasks. High-quality Text-to-video (T2V), a task that has been long considered mission-impossible, was proven feasible with…

Artificial Intelligence · Computer Science 2022-11-28 Yuxing Qiu , Feng Gao , Minchen Li , Govind Thattai , Yin Yang , Chenfanfu Jiang

We present a system using Multimodal LLMs (MLLMs) to analyze a large database with tens of millions of images captured at different times, with the aim of discovering patterns in temporal changes. Specifically, we aim to capture frequent…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Boyang Deng , Songyou Peng , Kyle Genova , Gordon Wetzstein , Noah Snavely , Leonidas Guibas , Thomas Funkhouser

Common-sense physical reasoning in the real world requires learning about the interactions of objects and their dynamics. The notion of an abstract object, however, encompasses a wide variety of physical objects that differ greatly in terms…

Machine Learning · Computer Science 2020-12-16 Aleksandar Stanić , Sjoerd van Steenkiste , Jürgen Schmidhuber

One of the fundamental goals of visual perception is to allow agents to meaningfully interact with their environment. In this paper, we take a step towards that long-term goal -- we extract highly localized actionable information related to…

Computer Vision and Pattern Recognition · Computer Science 2021-08-12 Kaichun Mo , Leonidas Guibas , Mustafa Mukadam , Abhinav Gupta , Shubham Tulsiani

Large vision-language models (VLMs) often benefit from intermediate visual cues, either injected via external tools or generated as latent visual tokens during reasoning, but these mechanisms still overlook fine-grained visual evidence…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Shuoshuo Zhang , Yizhen Zhang , Jingjing Fu , Lei Song , Jiang Bian , Yujiu Yang , Rui Wang

While current vision algorithms excel at many challenging tasks, it is unclear how well they understand the physical dynamics of real-world environments. Here we introduce Physion, a dataset and benchmark for rigorously evaluating the…

Scientists have long aimed to discover meaningful formulae which accurately describe experimental data. A common approach is to manually create mathematical models of natural phenomena using domain knowledge, and then fit these models to…

Understanding how the human brain represents visual concepts, and in which brain regions these representations are encoded, remains a long-standing challenge. Decades of work have advanced our understanding of visual representations, yet…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Navve Wasserman , Matias Cosarinsky , Yuval Golbari , Aude Oliva , Antonio Torralba , Tamar Rott Shaham , Michal Irani

Distilling knowledge from human demonstrations is a promising way for robots to learn and act. Existing methods, which often rely on coarsely-aligned video pairs, are typically constrained to learning global or task-level features. As a…

Robotics · Computer Science 2025-11-18 Sicheng Xie , Haidong Cao , Zejia Weng , Zhen Xing , Haoran Chen , Shiwei Shen , Jiaqi Leng , Zuxuan Wu , Yu-Gang Jiang

As Physics did in previous centuries, there is currently a common dream of extracting generic laws of nature in economics, sociology, neuroscience, by focalising the description of phenomena to a minimal set of variables and parameters,…

Physics and Society · Physics 2016-10-14 Fatihcan M. Atay , Sven Banisch , Philippe Blanchard , Bruno Cessac , Eckehard Olbrich