中文
相关论文

相关论文: RecipeGen: A Step-Aligned Multimodal Benchmark for…

200 篇论文

Recent advancements in text-to-image (T2I) generation models have transformed the field. However, challenges persist in generating images that reflect demanding textual descriptions, especially for fine-grained details and unusual…

多媒体 · 计算机科学 2025-02-21 Ran Li , Xiaomeng Jin , Heng ji

Text-and-Image-To-Image (TI2I), an extension of Text-To-Image (T2I), integrates image inputs with textual instructions to enhance image generation. Existing methods often partially utilize image inputs, focusing on specific elements like…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Teng-Fang Hsiao , Bo-Kai Ruan , Yi-Lun Wu , Tzu-Ling Lin , Hong-Han Shuai

Nutrition estimation is an important component of promoting healthy eating and mitigating diet-related health risks. Despite advances in tasks such as food classification and ingredient recognition, progress in nutrition estimation is…

计算机视觉与模式识别 · 计算机科学 2025-05-14 Huiyan Qi , Bin Zhu , Chong-Wah Ngo , Jingjing Chen , Ee-Peng Lim

Research on text-to-image generation has witnessed significant progress in generating diverse and photo-realistic images, driven by diffusion and auto-regressive models trained on large-scale image-text data. Though state-of-the-art models…

计算机视觉与模式识别 · 计算机科学 2022-11-23 Wenhu Chen , Hexiang Hu , Chitwan Saharia , William W. Cohen

Recent advances in text-to-video (T2V) generation highlight the critical role of high-quality video-text pairs in training models capable of producing coherent and instruction-aligned videos. However, strategies for optimizing video…

计算机视觉与模式识别 · 计算机科学 2025-10-31 Yang Du , Zhuoran Lin , Kaiqiang Song , Biao Wang , Zhicheng Zheng , Tiezheng Ge , Bo Zheng , Qin Jin

Evaluating the quality of videos generated from text-to-video (T2V) models is important if they are to produce plausible outputs that convince a viewer of their authenticity. We examine some of the metrics used in this area and highlight…

计算机视觉与模式识别 · 计算机科学 2023-09-18 Iya Chivileva , Philip Lynch , Tomas E. Ward , Alan F. Smeaton

To replicate the success of text-to-image (T2I) generation, recent works employ large-scale video datasets to train a text-to-video (T2V) generator. Despite their promising results, such paradigm is computationally expensive. In this work,…

计算机视觉与模式识别 · 计算机科学 2023-03-20 Jay Zhangjie Wu , Yixiao Ge , Xintao Wang , Weixian Lei , Yuchao Gu , Yufei Shi , Wynne Hsu , Ying Shan , Xiaohu Qie , Mike Zheng Shou

Text-to-image (T2I) generation aims to synthesize images from textual prompts, which jointly specify what must be shown and imply what can be inferred, which thus correspond to two core capabilities: \textbf{\textit{composition}} and…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Ouxiang Li , Yuan Wang , Xinting Hu , Huijuan Huang , Rui Chen , Jiarong Ou , Xin Tao , Pengfei Wan , Xiaojuan Qi , Fuli Feng

A rapidly growing amount of content posted online, such as food recipes, opens doors to new exciting applications at the intersection of vision and language. In this work, we aim to estimate the calorie amount of a meal directly from an…

计算机视觉与模式识别 · 计算机科学 2020-11-03 Robin Ruede , Verena Heusser , Lukas Frank , Alina Roitberg , Monica Haurilet , Rainer Stiefelhagen

Text-to-image (T2I) generation models have significantly advanced in recent years. However, effective interaction with these models is challenging for average users due to the need for specialized prompt engineering knowledge and the…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Minbin Huang , Yanxin Long , Xinchi Deng , Ruihang Chu , Jiangfeng Xiong , Xiaodan Liang , Hong Cheng , Qinglin Lu , Wei Liu

In this paper, we present an empirical study introducing a nuanced evaluation framework for text-to-image (T2I) generative models, applied to human image synthesis. Our framework categorizes evaluations into two distinct groups: first,…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Muxi Chen , Yi Liu , Jian Yi , Changran Xu , Qiuxia Lai , Hongliang Wang , Tsung-Yi Ho , Qiang Xu

Although recent text-to-image generative models have achieved impressive performance, they still often struggle with capturing the compositional complexities of prompts including attribute binding, and spatial relationships between…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Seyed Mohammad Hadi Hosseini , Amir Mohammad Izadi , Ali Abdollahi , Armin Saghafian , Mahdieh Soleymani Baghshah

Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text-video alignment, or physical…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Xianjing Han , Bin Zhu , Shiqi Hu , Franklin Mingzhe Li , Patrick Carrington , Roger Zimmermann , Jingjing Chen

Recent great advances in video generation models have demonstrated their potential to produce high-quality videos, bringing challenges to effective evaluation. Unlike human evaluation, existing automated evaluation metrics lack highlevel…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Zhun Mou , Bin Xia , Zhengchao Huang , Wenming Yang , Jiaya Jia

Diffusion-based text-to-video generation has witnessed impressive progress in the past year yet still falls behind text-to-image generation. One of the key reasons is the limited scale of publicly available data (e.g., 10M video-text pairs…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Xiang Wang , Shiwei Zhang , Hangjie Yuan , Zhiwu Qing , Biao Gong , Yingya Zhang , Yujun Shen , Changxin Gao , Nong Sang

Understanding food recipe requires anticipating the implicit causal effects of cooking actions, such that the recipe can be converted into a graph describing the temporal workflow of the recipe. This is a non-trivial task that involves…

计算与语言 · 计算机科学 2020-08-24 Liangming Pan , Jingjing Chen , Jianlong Wu , Shaoteng Liu , Chong-Wah Ngo , Min-Yen Kan , Yu-Gang Jiang , Tat-Seng Chua

Recent text-to-image (T2I) generation models have achieved remarkable sucess by training on billion-scale datasets, following a `bigger is better' paradigm that prioritizes data quantity over availability (closed vs open source) and…

计算机视觉与模式识别 · 计算机科学 2025-10-03 L. Degeorge , A. Ghosh , N. Dufour , D. Picard , V. Kalogeiton

Recent advances in text-to-video generation have produced increasingly realistic and diverse content, yet evaluating such videos remains a fundamental challenge due to their multi-faceted nature encompassing visual quality, semantic…

As Vision-Language Models (VLMs) increasingly gain traction in medical applications, clinicians are progressively expecting AI systems not only to generate textual diagnoses but also to produce corresponding medical images that integrate…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Junjie Yang , Yuhao Yan , Gang Wu , Yuxuan Wang , Ruoyu Liang , Xinjie Jiang , Xiang Wan , Fenglei Fan , Yongquan Zhang , Feiwei Qin , Changmiao Wang

Existing text-to-video (T2V) evaluation benchmarks, such as VBench and EvalCrafter, suffer from two limitations. (i) While the emphasis is on subject-centric prompts or static camera scenes, camera motion essential for producing cinematic…

计算机视觉与模式识别 · 计算机科学 2025-10-10 Nithin C. Babu , Aniruddha Mahapatra , Harsh Rangwani , Rajiv Soundararajan , Kuldeep Kulkarni