中文
相关论文

相关论文: Physion-Eval: Evaluating Physical Realism in Gener…

200 篇论文

Recent progress in generative video models, such as Veo-3, has shown surprising zero-shot reasoning abilities, creating a growing need for systematic and reliable evaluation. We introduce V-ReasonBench, a benchmark designed to assess video…

计算机视觉与模式识别 · 计算机科学 2025-11-21 Yang Luo , Xuanlei Zhao , Baijiong Lin , Lingting Zhu , Liyao Tang , Yuqi Liu , Ying-Cong Chen , Shengju Qian , Xin Wang , Yang You

Collecting real-world vehicle accident videos for autonomous driving research is challenging due to their rarity and complexity. While existing driving video generation methods may produce visually realistic videos, they often fail to…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Xiangwen Zhang , Qian Zhang , Longfei Han , Qiang Qu , Xiaoming Chen , Weidong Cai

While modern visual generation models excel at creating aesthetically pleasing natural images, they struggle with producing or editing structured visuals like charts, diagrams, and mathematical figures, which demand composition planning,…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Le Zhuo , Songhao Han , Yuandong Pu , Boxiang Qiu , Sayak Paul , Yue Liao , Yihao Liu , Jie Shao , Xi Chen , Si Liu , Hongsheng Li

Social interactions incorporate nonverbal signals to convey emotions alongside speech, including facial expressions and body gestures. Generative models have demonstrated promising results in creating full-body nonverbal animations…

人机交互 · 计算机科学 2026-04-01 Kiran Chhatre , Renan Guarese , Andrii Matviienko , Christopher Peters

Recent advancements in audio-visual generative modeling have been propelled by progress in deep learning and the availability of data-rich benchmarks. However, the growth is not attributed solely to models and benchmarks. Universally…

计算机视觉与模式识别 · 计算机科学 2024-04-12 Lucas Goncalves , Prashant Mathur , Chandrashekhar Lavania , Metehan Cekic , Marcello Federico , Kyu J. Han

Autonomous driving requires robust perception models trained on high-quality, large-scale multi-view driving videos for tasks like 3D object detection, segmentation and trajectory prediction. While world models provide a cost-effective…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Zhuoran Yang , Xi Guo , Chenjing Ding , Chiyu Wang , Wei Wu

Synthetic videos nowadays is widely used to complement data scarcity and diversity of real-world videos. Current synthetic datasets primarily replicate real-world scenarios, leaving impossible, counterfactual and anti-reality video concepts…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Zechen Bai , Hai Ci , Mike Zheng Shou

Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire day or more of continuous visual experience. In ultra-long video settings, relevant information is…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Ziyang Wang , Yue Zhang , Shoubin Yu , Ce Zhang , Zengqi Zhao , Jaehong Yoon , Hyunji Lee , Gedas Bertasius , Mohit Bansal

Emotion plays a pivotal role in video-based expression, but existing video generation systems predominantly focus on low-level visual metrics while neglecting affective dimensions. Although emotion analysis has made progress in the visual…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Zongyang Qiu , Bingyuan Wang , Xingbei Chen , Yingqing He , Zeyu Wang

Current video generation models produce high-quality aesthetic videos but often struggle to learn representations of real-world physics dynamics, resulting in artifacts such as unnatural object collisions, inconsistent gravity, and temporal…

计算机视觉与模式识别 · 计算机科学 2026-01-08 Siddarth Nilol Kundur Satish , Devesh Jaiswal , Hongyu Chen , Abhishek Bakshi

Despite recent progress in using Large Language Models (LLMs) for automatically generating 3D scenes, generated scenes often lack realistic spatial layouts and object attributes found in real-world environments. As this problem stems from…

计算与语言 · 计算机科学 2026-01-29 Gyeom Hwangbo , Hyungjoo Chae , Minseok Kang , Hyeonjong Ju , Soohyun Oh , Jinyoung Yeo

Recent advances in generative foundational models, often termed "world models," have propelled interest in applying them to critical tasks like robotic planning and autonomous system training. For reliable deployment, these models must…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Rishi Upadhyay , Howard Zhang , Jim Solomon , Ayush Agrawal , Pranay Boreddy , Shruti Satya Narayana , Yunhao Ba , Alex Wong , Celso M de Melo , Achuta Kadambi

Recent advances in deep generative models have led to significant progress in video generation, yet the fidelity of AI-generated videos remains limited. Synthesized content often exhibits visual artifacts such as temporally inconsistent…

计算机视觉与模式识别 · 计算机科学 2025-06-26 Jiahao Lin , Weixuan Peng , Bojia Zi , Yifeng Gao , Xianbiao Qi , Xingjun Ma , Yu-Gang Jiang

The recently developed Sora model [1] has exhibited remarkable capabilities in video generation, sparking intense discussions regarding its ability to simulate real-world phenomena. Despite its growing popularity, there is a lack of…

计算机视觉与模式识别 · 计算机科学 2024-02-28 Xuanyi Li , Daquan Zhou , Chenxu Zhang , Shaodong Wei , Qibin Hou , Ming-Ming Cheng

Generative world models hold significant potential for simulating interactions with visuomotor policies in varied environments. Frontier video models can enable generation of realistic observations and environment interactions in a scalable…

The capacity to create realistic virtual humans has progressed significantly, and such characters can be found in many applications across entertainment, education and health. As an essential element of interactive virtual humans,…

图形学 · 计算机科学 2026-05-12 Haoyang Du , Yinghan Xu , John Dingliana , Brian Keegan , Rachel McDonnell , Cathy Ennis

Foundation models have achieved remarkable success across video, image, and language domains. By scaling up the number of parameters and training datasets, these models acquire generalizable world knowledge and often surpass task-specific…

机器学习 · 计算机科学 2025-07-16 Tung Nguyen , Arsh Koneru , Shufan Li , Aditya Grover

Deep learning for human action recognition in videos is making significant progress, but is slowed down by its dependency on expensive manual labeling of large video collections. In this work, we investigate the generation of synthetic…

计算机视觉与模式识别 · 计算机科学 2017-07-20 César Roberto de Souza , Adrien Gaidon , Yohann Cabon , Antonio Manuel López Peña

Video generation models produce visually compelling results but systematically violate physical commonsense -- on VideoPhy-2, the best model achieves only 32.6% joint accuracy. We identify a specification bottleneck: text prompts are lossy…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Yuxiang Feng , Juncheng Wang , Chao Xu , Yijie Qian , Huihan Wang , Wenlong Hou , Yang Liu , Baigui Sun , Yong Liu , Shujun Wang

Building machines that can reason about physical events and their causal relationships is crucial for flexible interaction with the physical world. However, most existing physical and causal reasoning benchmarks are exclusively based on…

人工智能 · 计算机科学 2025-05-28 Jiayuan Mao , Xuelin Yang , Xikun Zhang , Noah D. Goodman , Jiajun Wu