中文
相关论文

相关论文: T2VPhysBench: A First-Principles Benchmark for Phy…

200 篇论文

World simulators can provide safe and scalable environments for training Physical AI systems before real-world deployment. Large video generation models are emerging as a promising basis for such simulators because they can generate diverse…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Pu Zhao , Juyi Lin , Timothy Rupprecht , Arash Akbari , Chence Yang , Rahul Chowdhury , Elaheh Motamedi , Arman Akbari , Yumei He , Chen Wang , Geng Yuan , Weiwei Chen , Yanzhi Wang

Despite the impressive advances in text-to-image models, they often struggle to effectively compose complex scenes with multiple objects, displaying various attributes and relationships. To address this challenge, we present…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Kaiyi Huang , Chengqi Duan , Kaiyue Sun , Enze Xie , Zhenguo Li , Xihui Liu

Text rendering has recently emerged as one of the most challenging frontiers in visual generation, drawing significant attention from large-scale diffusion and multimodal models. However, text editing within images remains largely…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Rui Gui , Yang Wan , Haochen Han , Dongxing Mao , Fangming Liu , Min Li , Alex Jinpeng Wang

Video generation models are increasingly used as world simulators for tasks like driving and robotic manipulation. What matters in these settings is not whether a single video looks right, but whether the model's output changes when its…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Kunlin Cai , Rui Song , Jinghuai Zhang , Kaiyuan Zhang , Pranav Bodapati , Alicia Yu , Fnu Suya , Mohammad Rostami , Jiaqi Ma , Yuan Tian

Generative video models achieve high visual fidelity but often violate basic physical principles, limiting reliability in real-world settings. Prior attempts to inject physics rely on conditioning: frame-level signals are domain-specific…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Saurabh Pathak , Elahe Arani , Mykola Pechenizkiy , Bahram Zonooz

With AI-generated videos increasingly indistinguishable from reality, current benchmarks primarily focus on broad semantic alignment and basic physical consistency, offering limited discriminative power for evaluating them. To address this,…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Jiaqi Wang , Weijia Wu , Yi Zhan , Rui Zhao , Ming Hu , James Cheng , Wei Liu , Philip Torr , Kevin Qinghong Lin

Despite advancements in generating visually stunning content, video diffusion models (VDMs) often yield physically inconsistent results due to pixel-only reconstruction. To address this, we propose MMPhysVideo, the first framework to scale…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Shubo Lin , Xuanyang Zhang , Wei Cheng , Weiming Hu , Gang Yu , Jin Gao

Continual post-training adapts a single text-to-image diffusion model to learn new tasks without incurring the cost of separate models, but naive post-training causes forgetting of pretrained knowledge and undermines zero-shot…

计算机视觉与模式识别 · 计算机科学 2025-05-23 Zhehao Huang , Yuhang Liu , Yixin Lou , Zhengbao He , Mingzhen He , Wenxing Zhou , Tao Li , Kehan Li , Zeyi Huang , Xiaolin Huang

This work aims to learn a high-quality text-to-video (T2V) generative model by leveraging a pre-trained text-to-image (T2I) model as a basis. It is a highly desirable yet challenging task to simultaneously a) accomplish the synthesis of…

A core capability towards general embodied intelligence lies in localizing task-relevant objects from an egocentric perspective, formulated as Spatio-Temporal Video Grounding (STVG). Despite recent progress, existing STVG studies remain…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Qi'ao Xu , Tianwen Qian , Yuqian Fu , Kailing Li , Yang Jiao , Jiacheng Zhang , Xiaoling Wang , Liang He

Over the past few years, Text-to-Image (T2I) generation approaches based on diffusion models have gained significant attention. However, vanilla diffusion models often suffer from spelling inaccuracies in the text displayed within the…

计算机视觉与模式识别 · 计算机科学 2024-10-30 Sanyam Lakhanpal , Shivang Chopra , Vinija Jain , Aman Chadha , Man Luo

Advances in technology have led to the development of methods that can create desired visual multimedia. In particular, image generation using deep learning has been extensively studied across diverse fields. In comparison, video…

计算机视觉与模式识别 · 计算机科学 2021-06-29 Doyeon Kim , Donggyu Joo , Junmo Kim

Existing benchmarks for assessing the spatio-temporal understanding and reasoning abilities of video language models are susceptible to score inflation due to the presence of shortcut solutions based on superficial visual or textual cues.…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Benno Krojer , Mojtaba Komeili , Candace Ross , Quentin Garrido , Koustuv Sinha , Nicolas Ballas , Mahmoud Assran

Large Vision-Language Models (LVLMs) have made significant strides in the field of video understanding in recent times. Nevertheless, existing video benchmarks predominantly rely on text prompts for evaluation, which often require complex…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Yiming Zhao , Yu Zeng , Yukun Qi , YaoYang Liu , Xikun Bao , Lin Chen , Zehui Chen , Qing Miao , Chenxi Liu , Jie Zhao , Feng Zhao

We introduce CausalVQA, a benchmark dataset for video question answering (VQA) composed of question-answer pairs that probe models' understanding of causality in the physical world. Existing VQA benchmarks either tend to focus on surface…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Aaron Foss , Chloe Evans , Sasha Mitts , Koustuv Sinha , Ammar Rizvi , Justine T. Kao

Recent advances in generative modeling can create remarkably realistic synthetic videos, making it increasingly difficult for humans to distinguish them from real ones and necessitating reliable detection methods. However, two key…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Long Ma , Zihao Xue , Yan Wang , Zhiyuan Yan , Jin Xu , Xiaorui Jiang , Haiyang Yu , Yong Liao , Zhen Bi

Video prediction is increasingly viewed as a path toward generalizable world models, yet it remains unclear whether these systems learn underlying causal structure or merely exploit superficial visual correlations for future prediction. We…

计算机视觉与模式识别 · 计算机科学 2026-05-25 León Begiristain , Olaf Dünkel , Adam Kortylewski

Existing text-to-video (T2V) models often struggle with generating videos with sufficiently pronounced or complex actions. A key limitation lies in the text prompt's inability to precisely convey intricate motion details. To address this,…

计算机视觉与模式识别 · 计算机科学 2024-11-14 Qiang Zhou , Shaofeng Zhang , Nianzu Yang , Ye Qian , Hao Li

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

机器学习 · 计算机科学 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

Recent rapid advancements in text-to-video (T2V) generation, such as SoRA and Kling, have shown great potential for building world simulators. However, current T2V models struggle to grasp abstract physical principles and generate videos…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Jing Wang , Ao Ma , Ke Cao , Jun Zheng , Zhanjie Zhang , Jiasong Feng , Shanyuan Liu , Yuhang Ma , Bo Cheng , Dawei Leng , Yuhui Yin , Xiaodan Liang
‹ 上一页 1 8 9 10 下一页 ›