中文
相关论文

相关论文: Towards World Simulator: Crafting Physical Commons…

200 篇论文

Visual Commonsense Reasoning, which is regarded as one challenging task to pursue advanced visual scene comprehension, has been used to diagnose the reasoning ability of AI systems. However, reliable reasoning requires a good grasp of the…

计算机视觉与模式识别 · 计算机科学 2025-01-17 Fan Yuan , Xiaoyuan Fang , Rong Quan , Jing Li , Wei Bi , Xiaogang Xu , Piji Li

Video generation assessment is essential for ensuring that generative models produce visually realistic, high-quality videos while aligning with human expectations. Current video generation benchmarks fall into two main categories:…

计算机视觉与模式识别 · 计算机科学 2025-04-30 Hui Han , Siyuan Li , Jiaqi Chen , Yiwen Yuan , Yuling Wu , Chak Tou Leong , Hanwen Du , Junchen Fu , Youhua Li , Jie Zhang , Chi Zhang , Li-jia Li , Yongxin Ni

Current Visual Language Models (VLMs) show impressive image understanding but struggle with visual illusions, especially in real-world scenarios. Existing benchmarks focus on classical cognitive illusions, which have been learned by…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Yiming Zhang , Zicheng Zhang , Xinyi Wei , Xiaohong Liu , Guangtao Zhai , Xiongkuo Min

Despite progress in video large language models (Video-LLMs), research on instructional video understanding, crucial for enhancing access to instructional content, remains insufficient. To address this, we introduce InstructionBench, an…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Haiwan Wei , Yitian Yuan , Xiaohan Lan , Wei Ke , Lin Ma

Recent advances in Text-to-3D (T23D) generative models have enabled the synthesis of diverse, high-fidelity 3D assets from textual prompts. However, existing challenges restrict the development of reliable T23D quality assessment (T23DQA).…

计算机视觉与模式识别 · 计算机科学 2025-11-06 Bingyang Cui , Yujie Zhang , Qi Yang , Zhu Li , Yiling Xu

Conversational generative vision models (CGVMs) like Visual ChatGPT (Wu et al., 2023) have recently emerged from the synthesis of computer vision and natural language processing techniques. These models enable more natural and interactive…

计算机视觉与模式识别 · 计算机科学 2023-05-30 Narjes Nikzad Khasmakhi , Meysam Asgari-Chenaghlu , Nabiha Asghar , Philipp Schaer , Dietlind Zühlke

Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text-video alignment, or physical…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Xianjing Han , Bin Zhu , Shiqi Hu , Franklin Mingzhe Li , Patrick Carrington , Roger Zimmermann , Jingjing Chen

The rapid advancements of Text-to-Image (T2I) models have ushered in a new phase of AI-generated content, marked by their growing ability to interpret and follow user instructions. However, existing T2I model evaluation benchmarks fall…

计算机视觉与模式识别 · 计算机科学 2025-06-26 Xinyu Wei , Jinrui Zhang , Zeqing Wang , Hongyang Wei , Zhen Guo , Lei Zhang

Video foundation models generate visually realistic and temporally coherent content, but their reliability as world simulators depends on whether they capture physical, logical, and spatial constraints. Existing metrics such as Frechet…

Recent advancements in predictive models have demonstrated exceptional capabilities in predicting the future state of objects and scenes. However, the lack of categorization based on inherent characteristics continues to hinder the progress…

计算机视觉与模式识别 · 计算机科学 2024-10-24 Yiran Qin , Zhelun Shi , Jiwen Yu , Xijun Wang , Enshen Zhou , Lijun Li , Zhenfei Yin , Xihui Liu , Lu Sheng , Jing Shao , Lei Bai , Wanli Ouyang , Ruimao Zhang

Understanding what constitutes safe text is an important issue in natural language processing and can often prevent the deployment of models deemed harmful and unsafe. One such type of safety that has been scarcely studied is commonsense…

计算与语言 · 计算机科学 2022-10-19 Sharon Levy , Emily Allaway , Melanie Subbiah , Lydia Chilton , Desmond Patton , Kathleen McKeown , William Yang Wang

Evaluating the quality of synthesized images remains a significant challenge in the development of text-to-image (T2I) generation. Most existing studies in this area primarily focus on evaluating text-image alignment, image quality, and…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Ziwei Huang , Wanggui He , Quanyu Long , Yandi Wang , Haoyuan Li , Zhelun Yu , Fangxun Shu , Long Chan , Hao Jiang , Fei Wu , Leilei Gan

Infographics are composite visual artifacts that combine data visualizations with textual and illustrative elements to communicate information. While recent text-to-image (T2I) models can generate aesthetically appealing images, their…

Despite rapid progress in multimodal large language models (MLLMs) and emerging omni-modal architectures, current benchmarks remain limited in scope and integration, suffering from incomplete modality coverage, restricted interaction to…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Yue Jiang , Dingkang Yang , Minghao Han , Jinghang Han , Zizhi Chen , Yizhou Liu , Mingcheng Li , Peng Zhai , Lihua Zhang

Multimodal Large Language Models (MLLMs) have achieved remarkable success in open-vocabulary perceptual tasks, yet their ability to solve complex cognitive problems remains limited, especially when visual details are abstract and require…

The emergence of unified multimodal understanding and generation models is rapidly attracting attention because of their ability to enhance instruction-following capabilities while minimizing model redundancy. However, there is a lack of a…

计算机视觉与模式识别 · 计算机科学 2025-05-16 Yi Li , Haonan Wang , Qixiang Zhang , Boyu Xiao , Chenchang Hu , Hualiang Wang , Xiaomeng Li

Interactive world models that simulate object dynamics are crucial for robotics, VR, and AR. However, it remains a significant challenge to learn physics-consistent dynamics models from limited real-world video data, especially for…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Yu Yang , Zhilu Zhang , Xiang Zhang , Yihan Zeng , Hui Li , Wangmeng Zuo

Recent advancements in text-to-video models such as Sora, Gen-3, MovieGen, and CogVideoX are pushing the boundaries of synthetic video generation, with adoption seen in fields like robotics, autonomous driving, and entertainment. As these…

计算机视觉与模式识别 · 计算机科学 2025-04-28 S P Sharan , Minkyu Choi , Sahil Shah , Harsh Goel , Mohammad Omama , Sandeep Chinchali

Surprising videos, such as funny clips, creative performances, or visual illusions, attract significant attention. Enjoyment of these videos is not simply a response to visual stimuli; rather, it hinges on the human capacity to understand…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Binzhu Xie , Sicheng Zhang , Zitang Zhou , Bo Li , Yuanhan Zhang , Jack Hessel , Jingkang Yang , Ziwei Liu

Recently, large-scale pre-trained language models have demonstrated impressive performance on several commonsense-reasoning benchmark datasets. However, building machines with commonsense to compose realistically plausible sentences remains…

计算与语言 · 计算机科学 2020-12-01 Bill Yuchen Lin , Wangchunshu Zhou , Ming Shen , Pei Zhou , Chandra Bhagavatula , Yejin Choi , Xiang Ren