中文
相关论文

相关论文: UniPPTBench: A Unified Benchmark for Presentation …

200 篇论文

As LLM agents are increasingly built around reusable skills, a central challenge is no longer only whether agents can use provided skills, but whether they can generate correct, reusable, and executable skills from repositories and…

Audio generation has long been fragmented, with speech, music, and sound effects produced by domain-specific models that fail to jointly generate coherent audio scenes from a single description. The key obstacles are insufficient…

Large Language Models (LLMs), with their exceptional ability to handle a wide range of tasks, have driven significant advancements in tackling reasoning and planning tasks, wherein decomposing complex problems into executable workflows is a…

计算与语言 · 计算机科学 2025-02-25 Shuofei Qiao , Runnan Fang , Zhisong Qiu , Xiaobin Wang , Ningyu Zhang , Yong Jiang , Pengjun Xie , Fei Huang , Huajun Chen

Story visualization aims to generate coherent image sequences that faithfully represent a narrative and match given character references. Despite progress in generative models, existing benchmarks remain narrow in scope, often limited to…

Interleaved text-and-image generation has been an intriguing research direction, where the models are required to generate both images and text pieces in an arbitrary order. Despite the emerging advancements in interleaved generation, the…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Minqian Liu , Zhiyang Xu , Zihao Lin , Trevor Ashby , Joy Rimchala , Jiaxin Zhang , Lifu Huang

While recent advancements in multimodal language models have enabled image generation from expressive multi-image instructions, existing methods struggle to maintain performance under complex interleaved instructions. This limitation stems…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Yabo Zhang , Kunchang Li , Dewei Zhou , Xinyu Huang , Xun Wang

In the scenario-based evaluation of machine learning models, a key problem is how to construct test datasets that represent various scenarios. The methodology proposed in this paper is to construct a benchmark and attach metadata to each…

软件工程 · 计算机科学 2024-06-19 Debalina Ghosh Paul , Hong Zhu , Ian Bayley

Recent image generation models excel at creating high-quality images from brief captions. However, they fail to maintain consistency of multiple instances across images when encountering lengthy contexts. This inconsistency is largely due…

计算机视觉与模式识别 · 计算机科学 2024-08-08 Zilyu Ye , Jinxiu Liu , Ruotian Peng , Jinjin Cao , Zhiyang Chen , Yiyang Zhang , Ziwei Xuan , Mingyuan Zhou , Xiaoqian Shen , Mohamed Elhoseiny , Qi Liu , Guo-Jun Qi

World Generation Models are emerging as a cornerstone of next-generation multimodal intelligence systems. Unlike traditional 2D visual generation, World Models aim to construct realistic, dynamic, and physically consistent 3D/4D worlds from…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Yiting Lu , Wei Luo , Peiyan Tu , Haoran Li , Hanxin Zhu , Zihao Yu , Xingrui Wang , Xinyi Chen , Xinge Peng , Xin Li , Zhibo Chen

Generative diffusion models are developing rapidly and attracting increasing attention due to their wide range of applications. Image-to-Video (I2V) generation has become a major focus in the field of video synthesis. However, existing…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Ailing Zhang , Lina Lei , Dehong Kong , Zhixin Wang , Jiaqi Xu , Fenglong Song , Chun-Le Guo , Chang Liu , Fan Li , Jie Chen

The next frontier for video generation lies in developing models capable of zero-shot reasoning, where understanding real-world scientific laws is crucial for accurate physical outcome modeling under diverse conditions. However, existing…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Lanxiang Hu , Abhilash Shankarampeta , Yixin Huang , Zilin Dai , Haoyang Yu , Yujie Zhao , Haoqiang Kang , Daniel Zhao , Tajana Rosing , Hao Zhang

Conditional time series generation plays a critical role in addressing data scarcity and enabling causal analysis in real-world applications. Despite its increasing importance, the field lacks a standardized and systematic benchmarking…

机器学习 · 计算机科学 2026-03-06 Shaocheng Lan , Shuqi Gu , Zhangzhi Xiong , Kan Ren

The evolution of video generation toward complex, multi-shot narratives has exposed a critical deficit in current evaluation methods. Existing benchmarks remain anchored to single-shot paradigms, lacking the comprehensive story assets and…

多媒体 · 计算机科学 2026-03-02 Haoyuan Shi , Yunxin Li , Nanhao Deng , Zhenran Xu , Xinyu Chen , Longyue Wang , Baotian Hu , Min Zhang

Multi-model learning has attracted great attention in visual-text tasks. However, visual-tabular data, which plays a pivotal role in high-stakes domains like healthcare and industry, remains underexplored. In this paper, we introduce…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Zi-Yi Jia , Zi-Jian Cheng , Xin-Yue Zhang , Kun-Yang Yu , Zhi Zhou , Yu-Feng Li , Lan-Zhe Guo

Recent years have witnessed a rapid development of deep generative models for creating synthetic media, such as images and videos. While the practical applications of these models in everyday tasks are enticing, it is crucial to assess the…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Mike Laszkiewicz , Imant Daunhawer , Julia E. Vogt , Asja Fischer , Johannes Lederer

Despite recent advances in text-conditioned 3D indoor scene generation, there remain gaps in the evaluation of these methods. Existing metrics often measure realism by comparing generated scenes to a set of ground-truth scenes, but they…

图形学 · 计算机科学 2026-03-10 Hou In Ivan Tam , Hou In Derek Pun , Austin T. Wang , Angel X. Chang , Manolis Savva

Video generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align…

Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to the input text and self-consistent across generated images…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Silin Gao , Sheryl Mathew , Li Mi , Sepideh Mamooler , Mengjie Zhao , Hiromi Wakaki , Yuki Mitsufuji , Syrielle Montariol , Antoine Bosselut

The evaluation of visual editing models remains fragmented across methods and modalities. Existing benchmarks are often tailored to specific paradigms, making fair cross-paradigm comparisons difficult, while video editing lacks reliable…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Lifan Jiang , Tianrun Wu , Yuhang Pei , Chenyang Wang , Boxi Wu , Deng Cai

Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models (Omni-LLMs), which are capable of processing multi-modal information including vision and audio,…

多媒体 · 计算机科学 2026-05-15 Jianghan Chao , Jianzhang Gao , Wenhui Tan , Yuchong Sun , Ruihua Song , Liyun Ru