中文
相关论文

相关论文: PhysInOne: Visual Physics Learning and Reasoning i…

200 篇论文

The autonomous evolution of networked AI systems relies heavily on robust environmental perception. However, physical understanding remains brittle in current models because key physical signals are visually ambiguous and sparsely…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Minghao Han , Dingkang Yang , Yue Jiang , Yizhou Liu , Lihua Zhang

Datasets advance research by posing challenging new problems and providing standardized methods of algorithm comparison. High-quality datasets exist for many important problems in robotics and computer vision, including egomotion estimation…

机器人学 · 计算机科学 2019-06-14 Kevin M. Judd , Jonathan D. Gammell

In this work, we propose a unified framework, called Visual Reasoning with Differ-entiable Physics (VRDP), that can jointly learn visual concepts and infer physics models of objects and their interactions from videos and language. This is…

计算机视觉与模式识别 · 计算机科学 2021-10-29 Mingyu Ding , Zhenfang Chen , Tao Du , Ping Luo , Joshua B. Tenenbaum , Chuang Gan

The increasing popularity of exercises including yoga and Pilates has created a greater demand for professional exercise video datasets in the realm of artificial intelligence. In this study, we developed 3DYoga901, which is organized…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Seonok Kim

We present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds. Modeling the spatial-audio-temporal dynamics even for actions…

计算机视觉与模式识别 · 计算机科学 2019-02-19 Mathew Monfort , Alex Andonian , Bolei Zhou , Kandan Ramakrishnan , Sarah Adel Bargal , Tom Yan , Lisa Brown , Quanfu Fan , Dan Gutfruend , Carl Vondrick , Aude Oliva

3D scene understanding is a long-standing challenge in computer vision and a key component in enabling mixed reality, wearable computing, and embodied AI. Providing a solution to these applications requires a multifaceted approach that…

计算机视觉与模式识别 · 计算机科学 2025-07-03 Anna-Maria Halacheva , Yang Miao , Jan-Nico Zaech , Xi Wang , Luc Van Gool , Danda Pani Paudel

Multiphysics phenomena, the coupling effects involving different aspects of physics laws, are pervasive in the real world and can often be encountered when performing everyday household tasks. Intelligent agents which seek to assist or…

机器人学 · 计算机科学 2023-05-16 Haoyuan Fu , Wenqiang Xu , Ruolin Ye , Han Xue , Zhenjun Yu , Tutian Tang , Yutong Li , Wenxin Du , Jieyi Zhang , Cewu Lu

Leveraging physical knowledge described by partial differential equations (PDEs) is an appealing way to improve unsupervised video prediction methods. Since physics is too restrictive for describing the full visual content of generic…

计算机视觉与模式识别 · 计算机科学 2020-03-18 Vincent Le Guen , Nicolas Thome

A deep understanding of the physical world is a central goal for embodied AI and realistic simulation. While current models excel at capturing an object's surface geometry and appearance, they largely neglect its internal physical…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Jingxuan Zhang , Tianqi Yu , Yatu Zhang , Jinze Wu , Kaixin Yao , Jingyang Liu , Yuyao Zhang , Jiayuan Gu , Jingyi Yu

Advances in neural fields are enabling high-fidelity capture of the shape and appearance of dynamic 3D scenes. However, their capabilities lag behind those offered by conventional representations such as 2D videos because of algorithmic…

The objective of this work is to develop an AI foundation model for physical signals that can generalize across diverse phenomena, domains, applications, and sensing apparatuses. We propose a phenomenological approach and framework for…

Estimating the geometric and volumetric properties of transparent deformable liquids is challenging due to optical complexities and dynamic surface deformations induced by container movements. Autonomous robots performing precise liquid…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Ke Ma , Yizhou Fang , Jean-Baptiste Weibel , Shuai Tan , Xinggang Wang , Yang Xiao , Yi Fang , Tian Xia

Accurate 6D object pose estimation from images is a key problem in object-centric scene understanding, enabling applications in robotics, augmented reality, and scene reconstruction. Despite recent advances, existing methods often produce…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Martin Malenický , Martin Cífka , Médéric Fourmy , Louis Montaut , Justin Carpentier , Josef Sivic , Vladimir Petrik

Physical video understanding requires more than naming an event correctly. A model can answer a question about pouring, sliding, or collision from textual regularities while still failing to localize the event in time or space. We introduce…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Alibay Osmanli , Zixu Cheng , Shaogang Gong

This document is a hands-on, comprehensive guide to deep learning in the realm of physical simulations. Rather than just theory, we emphasize practical application: every concept is paired with interactive Jupyter notebooks to get you up…

机器学习 · 计算机科学 2025-03-28 N. Thuerey , B. Holzschuh , P. Holl , G. Kohl , M. Lino , Q. Liu , P. Schnell , F. Trost

3D modeling is moving from virtual to physical. Existing 3D generation primarily emphasizes geometries and textures while neglecting physical-grounded modeling. Consequently, despite the rapid development of 3D generative models, the…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Ziang Cao , Zhaoxi Chen , Liang Pan , Ziwei Liu

Evaluating whether Multimodal Large Language Models (MLLMs) genuinely reason about physical dynamics remains challenging. Most existing benchmarks rely on recognition-style protocols such as Visual Question Answering (VQA) and Violation of…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Jiarong Liang , Max Ku , Ka-Hei Hui , Ping Nie , Wenhu Chen

We introduce an approach to model surface properties governing bounces in everyday scenes. Our model learns end-to-end, starting from sensor inputs, to predict post-bounce trajectories and infer two underlying physical properties that…

计算机视觉与模式识别 · 计算机科学 2019-04-16 Senthil Purushwalkam , Abhinav Gupta , Danny M. Kaufman , Bryan Russell

Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natural variations…

计算机视觉与模式识别 · 计算机科学 2022-05-03 Fadime Sener , Dibyadip Chatterjee , Daniel Shelepov , Kun He , Dipika Singhania , Robert Wang , Angela Yao

Despite significant advances in video generation, synthesizing physically plausible human actions remains a persistent challenge, particularly in modeling fine-grained semantics and complex temporal dynamics. For instance, generating…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Dian Shao , Mingfei Shi , Shengda Xu , Haodong Chen , Yongle Huang , Binglu Wang