中文
相关论文

相关论文: PhysInOne: Visual Physics Learning and Reasoning i…

200 篇论文

In recent years, deep neural network approaches have naturally extended to the video domain, in their simplest case by aggregating per-frame classifications as a baseline for action recognition. A majority of the work in this area extends…

计算机视觉与模式识别 · 计算机科学 2018-01-24 Daniel Castro , Steven Hickson , Patsorn Sangkloy , Bhavishya Mittal , Sean Dai , James Hays , Irfan Essa

Analyzing student behavior in educational scenarios is crucial for enhancing teaching quality and student engagement. Existing AI-based models often rely on classroom video footage to identify and analyze student behavior. While these…

计算机与社会 · 计算机科学 2025-03-11 Xian Gao , Jiacheng Ruan , Jingsheng Gao , Mingye Xie , Zongyun Zhang , Ting Liu , Yuzhuo Fu

Vision-Language Models (VLMs) have demonstrated strong performance on textbook-style physics problems, yet they frequently fail when confronted with dynamic real-world scenarios that require temporal consistency and causal reasoning across…

人工智能 · 计算机科学 2026-04-28 Sinin Zhang , Yunfei Xie , Yuxuan Cheng , Haoyu Zhang , Tong Zhang

Vision-language-action models have advanced rapidly, but robot trajectories alone provide limited coverage for learning broad physical understanding. PhysBrain 1.0 studies a complementary route: converting large-scale human egocentric video…

Recent rapid advancements in text-to-video (T2V) generation, such as SoRA and Kling, have shown great potential for building world simulators. However, current T2V models struggle to grasp abstract physical principles and generate videos…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Jing Wang , Ao Ma , Ke Cao , Jun Zheng , Zhanjie Zhang , Jiasong Feng , Shanyuan Liu , Yuhang Ma , Bo Cheng , Dawei Leng , Yuhui Yin , Xiaodan Liang

Large vision-language models exhibit inherent capabilities to handle diverse visual perception tasks. In this paper, we introduce VisionReasoner, a unified framework capable of reasoning and solving multiple visual perception tasks within a…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Yuqi Liu , Tianyuan Qu , Zhisheng Zhong , Bohao Peng , Shu Liu , Bei Yu , Jiaya Jia

Autonomous trucking is a promising technology that can greatly impact modern logistics and the environment. Ensuring its safety on public roads is one of the main duties that requires an accurate perception of the environment. To achieve…

Simulation-ready physical 3D assets have emerged as a promising direction owing to their broad applicability in downstream tasks. However, most existing 3D generation methods either neglect physical properties or are limited to a single…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Ziang Cao , Yinghao Liu , Haitian Li , Runmao Yao , Fangzhou Hong , Zhaoxi Chen , Liang Pan , Ziwei Liu

Single image scene relighting aims to generate a realistic new version of an input image so that it appears to be illuminated by a new target light condition. Although existing works have explored this problem from various perspectives,…

计算机视觉与模式识别 · 计算机科学 2024-09-30 Yixiong Yang , Hassan Ahmed Sial , Ramon Baldrich , Maria Vanrell

Dense and versatile image representations underpin the success of virtually all computer vision applications. However, state-of-the-art networks, such as transformers, produce low-resolution feature grids, which are suboptimal for dense…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Nikita Araslanov , Anna Sonnweber , Daniel Cremers

Dynamic environments such as urban areas are still challenging for popular visual-inertial odometry (VIO) algorithms. Existing datasets typically fail to capture the dynamic nature of these environments, therefore making it difficult to…

机器人学 · 计算机科学 2021-02-12 Koji Minoda , Fabian Schilling , Valentin Wüest , Dario Floreano , Takehisa Yairi

Video inpainting is the task of filling a region in a video in a visually convincing manner. It is very challenging due to the high dimensionality of the data and the temporal consistency required for obtaining convincing results. Recently,…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Nicolas Cherel , Andrés Almansa , Yann Gousseau , Alasdair Newson

The development of large-scale 3D scene reconstruction and novel view synthesis methods mostly rely on datasets comprising perspective images with narrow fields of view (FoV). While effective for small-scale scenes, these datasets require…

计算机视觉与模式识别 · 计算机科学 2025-04-10 Ulas Gunes , Matias Turkulainen , Xuqian Ren , Arno Solin , Juho Kannala , Esa Rahtu

We present WayveScenes101, a dataset designed to help the community advance the state of the art in novel view synthesis that focuses on challenging driving scenes containing many dynamic and deformable elements with changing geometry and…

计算机视觉与模式识别 · 计算机科学 2024-07-12 Jannik Zürn , Paul Gladkov , Sofía Dudas , Fergal Cotter , Sofi Toteva , Jamie Shotton , Vasiliki Simaiaki , Nikhil Mohan

Video generative models have made remarkable progress, yet they often yield visual artifacts that violate grounding in physical dynamics. Recent works such as PhysGen3D tackle single image-to-3D physics through mesh reconstruction and…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Hwidong Kim , Yunho Kim , Tae-Kyun Kim

Autonomous parking remains a critical yet challenging task in intelligent driving systems, particularly within constrained urban environments where maneuvering space is limited and precise control is essential. While recent advances in…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Haonan Chen , Kaiwen Xiao , Bin Tian , Jun Fu

From just a glance, humans can make rich predictions about the future state of a wide range of physical systems. On the other hand, modern approaches from engineering, robotics, and graphics are often restricted to narrow domains and…

计算机视觉与模式识别 · 计算机科学 2017-06-06 Nicholas Watters , Andrea Tacchetti , Theophane Weber , Razvan Pascanu , Peter Battaglia , Daniel Zoran

Visual augmentation has become a crucial technique for enhancing the visual robustness of imitation learning. However, existing methods are often limited by prerequisites such as camera calibration or the need for controlled environments…

机器人学 · 计算机科学 2025-07-15 Chengbo Yuan , Suraj Joshi , Shaoting Zhu , Hang Su , Hang Zhao , Yang Gao

Recent advances in video generation models demonstrate their potential as world simulators, but they often struggle with videos deviating from physical laws, a key concern overlooked by most text-to-video benchmarks. We introduce a…

计算机视觉与模式识别 · 计算机科学 2025-03-06 Yongfan Chen , Xiuwen Zhu , Tianyu Li

Most existing robotic datasets capture static scene data and thus are limited in evaluating robots' dynamic performance. To address this, we present a mobile robot oriented large-scale indoor dataset, denoted as THUD (Tsinghua University…

机器人学 · 计算机科学 2024-07-02 Yifan Tang , Cong Tai , Fangxing Chen , Wanting Zhang , Tao Zhang , Xueping Liu , Yongjin Liu , Long Zeng