中文
相关论文

相关论文: PhysVid: Physics Aware Local Conditioning for Gene…

200 篇论文

The existing state-of-the-art method for audio-visual conditioned video prediction uses the latent codes of the audio-visual frames from a multimodal stochastic network and a frame encoder to predict the next visual frame. However, a direct…

计算机视觉与模式识别 · 计算机科学 2023-09-21 Yating Xu , Conghui Hu , Gim Hee Lee

In recent years, there has been rapid development in 3D generation models, opening up new possibilities for applications such as simulating the dynamic movements of 3D objects and customizing their behaviors. However, current 3D generative…

计算机视觉与模式识别 · 计算机科学 2024-06-12 Fangfu Liu , Hanyang Wang , Shunyu Yao , Shengjun Zhang , Jie Zhou , Yueqi Duan

Vision models are often vulnerable to out-of-distribution (OOD) samples without adapting. While visual prompts offer a lightweight method of input-space adaptation for large-scale vision models, they rely on a high-dimensional additive…

计算机视觉与模式识别 · 计算机科学 2023-10-27 Yun-Yun Tsai , Chengzhi Mao , Junfeng Yang

Physical AI aims to develop models that can perceive and predict real-world dynamics; yet, the extent to which current multi-modal large language models and video generative models support these abilities is insufficiently understood. We…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Fengzhe Zhou , Jiannan Huang , Jialuo Li , Deva Ramanan , Humphrey Shi

Physical reasoning remains a significant challenge for Vision-Language Models (VLMs). This limitation arises from an inability to translate learned knowledge into predictions about physical behavior. Although continual fine-tuning can…

计算机视觉与模式识别 · 计算机科学 2025-10-30 Vahid Balazadeh , Mohammadmehdi Ataei , Hyunmin Cheong , Amir Hosein Khasahmadi , Rahul G. Krishnan

Creating a vivid video from the event or scenario in our imagination is a truly fascinating experience. Recent advancements in text-to-video synthesis have unveiled the potential to achieve this with prompts only. While text is convenient…

计算机视觉与模式识别 · 计算机科学 2023-06-02 Jinbo Xing , Menghan Xia , Yuxin Liu , Yuechen Zhang , Yong Zhang , Yingqing He , Hanyuan Liu , Haoxin Chen , Xiaodong Cun , Xintao Wang , Ying Shan , Tien-Tsin Wong

Generating controllable character animation from a reference image and motion guidance remains a challenging task due to the inherent difficulty of injecting appearance and motion cues into video diffusion models. Prior works often rely on…

图形学 · 计算机科学 2025-07-03 Guian Fang , Yuchao Gu , Mike Zheng Shou

We introduce PhysMotion, a novel framework that leverages principled physics-based simulations to guide intermediate 3D representations generated from a single image and input conditions (e.g., applied force and torque), producing…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Xiyang Tan , Ying Jiang , Xuan Li , Zeshun Zong , Tianyi Xie , Yin Yang , Chenfanfu Jiang

Synthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross-attention for text…

计算机视觉与模式识别 · 计算机科学 2025-08-18 Jiasong Feng , Ao Ma , Jing Wang , Ke Cao , Zhanjie Zhang

Current video generation models cannot simulate physical consequences of 3D actions like forces and robotic manipulations, as they lack structural understanding of how actions affect 3D scenes. We present RealWonder, the first real-time…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Wei Liu , Ziyu Chen , Zizhang Li , Yue Wang , Hong-Xing Yu , Jiajun Wu

Programmatic video generation through code offers geometric precision and temporal coherence beyond pixel-level diffusion models, yet rigorously evaluating whether language models can produce spatially correct animated outputs remains an…

人工智能 · 计算机科学 2026-05-20 Qiran Zhang , Yuheng Wang , Runde Yang , Lin Wu , Jingru Fan , Shu Yao , Jie Zhang , Tianle Zhou , Huatao Li , Ruijie Shi , Yihan Li , Chen Qian

By providing substantial amounts of data and standardized evaluation protocols, datasets in computer vision have helped fuel advances across all areas of visual recognition. But even in light of breakthrough results on recent benchmarks, it…

计算机视觉与模式识别 · 计算机科学 2018-07-06 Brandon RichardWebster , Samuel E. Anthony , Walter J. Scheirer

Recent advancements in video generation, particularly in diffusion models, have driven notable progress in text-to-video (T2V) and image-to-video (I2V) synthesis. However, challenges remain in effectively integrating dynamic motion signals…

计算机视觉与模式识别 · 计算机科学 2025-07-04 Ziye Li , Hao Luo , Xincheng Shuai , Henghui Ding

Current video models fail as world model as they lack fine-graiend control. General-purpose household robots require real-time fine motor control to handle delicate tasks and urgent situations. In this work, we introduce fine-grained…

计算机视觉与模式识别 · 计算机科学 2025-10-03 Yichen Li , Antonio Torralba

Learning a physical model from video data that can comprehend physical laws and predict the future trajectories of objects is a formidable challenge in artificial intelligence. Prior approaches either leverage various Partial Differential…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Nengbo Lu , Minghua Pan

We investigate how to enhance the physical fidelity of video generation models by leveraging synthetic videos derived from computer graphics pipelines. These rendered videos respect real-world physics, such as maintaining 3D consistency,…

图像与视频处理 · 电气工程与系统科学 2025-03-28 Qi Zhao , Xingyu Ni , Ziyu Wang , Feng Cheng , Ziyan Yang , Lu Jiang , Bohan Wang

A long-standing question in physical reasoning is whether video-based models need to rely on factorized representations of physical variables in order to make physically accurate predictions, or whether they can implicitly represent such…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Sonia Joseph , Quentin Garrido , Randall Balestriero , Matthew Kowal , Thomas Fel , Shahab Bakhtiari , Blake Richards , Mike Rabbat

Vision-Language Models (VLMs) have demonstrated strong performance on textbook-style physics problems, yet they frequently fail when confronted with dynamic real-world scenarios that require temporal consistency and causal reasoning across…

人工智能 · 计算机科学 2026-04-28 Sinin Zhang , Yunfei Xie , Yuxuan Cheng , Haoyu Zhang , Tong Zhang

The task of video grounding, which temporally localizes a natural language description in a video, plays an important role in understanding videos. Existing studies have adopted strategies of sliding window over the entire video or…

计算机视觉与模式识别 · 计算机科学 2019-01-23 Dongliang He , Xiang Zhao , Jizhou Huang , Fu Li , Xiao Liu , Shilei Wen

With the rapid development of large multimodal models, reliable judge and critic models have become essential for open-ended evaluation and preference alignment, providing pairwise preferences, numerical scores, and explanatory…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Tianyi Xiong , Shihao Wang , Guilin Liu , Yi Dong , Ming Li , Heng Huang , Jan Kautz , Zhiding Yu