中文
相关论文

相关论文: T2VPhysBench: A First-Principles Benchmark for Phy…

200 篇论文

Video large language models (Video-LLMs) have made strong progress in general video understanding, but their ability to maintain temporal object consistency remains underexplored. Existing benchmarks often emphasize event recognition,…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Junzhe Chen , Siyuan Meng , Yuxi Chen , Man Zhao , Wenyao Gui , Xiaojie Guo

Text-to-image (T2I) models have garnered significant attention for generating high-quality images aligned with text prompts. However, rapid T2I model advancements reveal limitations in early benchmarks, lacking comprehensive evaluations,…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Jingjing Chang , Yixiao Fang , Peng Xing , Shuhan Wu , Wei Cheng , Rui Wang , Xianfang Zeng , Gang Yu , Hai-Bao Chen

Video generation has achieved remarkable progress, with generated videos increasingly resembling real ones. However, the rapid advance in generation has outpaced the development of adequate evaluation metrics. Currently, the assessment of…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Nabyl Quignon , Baptiste Chopin , Yaohui Wang , Antitza Dantcheva

Reasoning is a fundamental capability often required in real-world text-to-image (T2I) generation, e.g., generating ``a bitten apple that has been left in the air for more than a week`` necessitates understanding temporal decay and…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Kaijie Chen , Zihao Lin , Zhiyang Xu , Ying Shen , Yuguang Yao , Joy Rimchala , Jiaxin Zhang , Lifu Huang

Subject-driven text-to-image (T2I) generation aims to produce images that align with a given textual description, while preserving the visual identity from a referenced subject image. Despite its broad downstream applicability - ranging…

Modern video diffusion models excel at appearance synthesis but still struggle with physical consistency: objects drift, collisions lack realistic rebound, and material responses seldom match their underlying properties. We present PhyCo, a…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Sriram Narayanan , Ziyu Jiang , Srinivasa Narasimhan , Manmohan Chandraker

Text-to-image (T2I) models today are capable of producing photorealistic, instruction-following images, yet they still frequently fail on prompts that require implicit world knowledge. Existing evaluation protocols either emphasize…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Tianyang Han , Junhao Su , Junjie Hu , Peizhen Yang , Hengyu Shi , Junfeng Luo , Jialin Gao

Text-to-image models are known to struggle with generating images that perfectly align with textual prompts. Several previous studies have focused on evaluating image-text alignment in text-to-image generation. However, these evaluations…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Huixuan Zhang , Xiaojun Wan

Image-to-video (I2V) generation seeks to produce realistic motion sequences from a single reference image. Although recent methods exhibit strong temporal consistency, they often struggle when dealing with complex, non-repetitive human…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Ashkan Taghipour , Morteza Ghahremani , Mohammed Bennamoun , Farid Boussaid , Aref Miri Rekavandi , Zinuo Li , Qiuhong Ke , Hamid Laga

The rapid advancements of Text-to-Image (T2I) models have ushered in a new phase of AI-generated content, marked by their growing ability to interpret and follow user instructions. However, existing T2I model evaluation benchmarks fall…

计算机视觉与模式识别 · 计算机科学 2025-06-26 Xinyu Wei , Jinrui Zhang , Zeqing Wang , Hongyang Wei , Zhen Guo , Lei Zhang

Foundation models have achieved remarkable success across video, image, and language domains. By scaling up the number of parameters and training datasets, these models acquire generalizable world knowledge and often surpass task-specific…

机器学习 · 计算机科学 2025-07-16 Tung Nguyen , Arsh Koneru , Shufan Li , Aditya Grover

Current text-to-video models (T2V) can generate high-quality, temporally coherent, and visually realistic videos. Nonetheless, errors still often occur, and are more nuanced and local compared to the previous generation of T2V models. While…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Aditya Chinchure , Sahithya Ravi , Pushkar Shukla , Vered Shwartz , Leonid Sigal

This paper introduces ModelScopeT2V, a text-to-video synthesis model that evolves from a text-to-image synthesis model (i.e., Stable Diffusion). ModelScopeT2V incorporates spatio-temporal blocks to ensure consistent frame generation and…

计算机视觉与模式识别 · 计算机科学 2023-08-15 Jiuniu Wang , Hangjie Yuan , Dayou Chen , Yingya Zhang , Xiang Wang , Shiwei Zhang

Evaluating text-to-image generative models remains a challenge, despite the remarkable progress being made in their overall performances. While existing metrics like CLIPScore work for coarse evaluations, they lack the sensitivity to…

计算机视觉与模式识别 · 计算机科学 2024-11-06 Georgia Gabriela Sampaio , Ruixiang Zhang , Shuangfei Zhai , Jiatao Gu , Josh Susskind , Navdeep Jaitly , Yizhe Zhang

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped…

计算机视觉与模式识别 · 计算机科学 2025-12-25 Zhe Cao , Tao Wang , Jiaming Wang , Yanghai Wang , Yuanxing Zhang , Jialu Chen , Miao Deng , Jiahao Wang , Yubin Guo , Chenxi Liao , Yize Zhang , Zhaoxiang Zhang , Jiaheng Liu

While Text-To-Video (T2V) models have advanced rapidly, they continue to struggle with generating legible and coherent text within videos. In particular, existing models often fail to render correctly even short phrases or words and…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Ziyang Liu , Kevin Valencia , Justin Cui

Text-to-image (T2I) generation aims to synthesize images from textual prompts, which jointly specify what must be shown and imply what can be inferred, which thus correspond to two core capabilities: \textbf{\textit{composition}} and…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Ouxiang Li , Yuan Wang , Xinting Hu , Huijuan Huang , Rui Chen , Jiarong Ou , Xin Tao , Pengfei Wan , Xiaojuan Qi , Fuli Feng

Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally ambiguous for fine-grained motion control, while explicit…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Ashkan Taghipour , Morteza Ghahremani , Zinuo Li , Hamid Laga , Farid Boussaid , Mohammed Bennamoun

Image-to-video (I2V) generation aims to use the initial frame (alongside a text prompt) to create a video sequence. A grand challenge in I2V generation is to maintain visual consistency throughout the video: existing methods often struggle…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Weiming Ren , Huan Yang , Ge Zhang , Cong Wei , Xinrun Du , Wenhao Huang , Wenhu Chen

Diffusion-based text-to-video generation has witnessed impressive progress in the past year yet still falls behind text-to-image generation. One of the key reasons is the limited scale of publicly available data (e.g., 10M video-text pairs…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Xiang Wang , Shiwei Zhang , Hangjie Yuan , Zhiwu Qing , Biao Gong , Yingya Zhang , Yujun Shen , Changxin Gao , Nong Sang