English
Related papers

Related papers: RecipeGen: A Step-Aligned Multimodal Benchmark for…

200 papers

Text-to-image (T2I) models, such as Stable Diffusion, have exhibited remarkable performance in generating high-quality images from text descriptions in recent years. However, text-to-image models may be tricked into generating…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Xinfeng Li , Yuchen Yang , Jiangyi Deng , Chen Yan , Yanjiao Chen , Xiaoyu Ji , Wenyuan Xu

Recently, Text-to-Image (T2I) generation models have achieved significant advancements. Correspondingly, many automated metrics have emerged to evaluate the image-text alignment capabilities of generative models. However, the performance…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Shuhao Han , Haotian Fan , Jiachen Fu , Liang Li , Tao Li , Junhui Cui , Yunqiu Wang , Yang Tai , Jingwei Sun , Chunle Guo , Chongyi Li

Recent advancements in camera-trajectory-guided image-to-video generation offer higher precision and better support for complex camera control compared to text-based approaches. However, they also introduce significant usability challenges,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Teng Li , Guangcong Zheng , Rui Jiang , Shuigen Zhan , Tao Wu , Yehao Lu , Yining Lin , Chuanyun Deng , Yepan Xiong , Min Chen , Lin Cheng , Xi Li

In recent years, Text-to-Image (T2I) models have been extensively studied, especially with the emergence of diffusion models that achieve state-of-the-art results on T2I synthesis tasks. However, existing benchmarks heavily rely on…

Computer Vision and Pattern Recognition · Computer Science 2023-11-27 Eslam Mohamed Bakr , Pengzhan Sun , Xiaoqian Shen , Faizan Farooq Khan , Li Erran Li , Mohamed Elhoseiny

Video generation models have achieved remarkable progress in creating high-quality, photorealistic content. However, their ability to accurately simulate physical phenomena remains a critical and unresolved challenge. This paper presents…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Jing Gu , Xian Liu , Yu Zeng , Ashwin Nagarajan , Fangrui Zhu , Daniel Hong , Yue Fan , Qianqi Yan , Kaiwen Zhou , Ming-Yu Liu , Xin Eric Wang

This paper reports on the NTIRE 2025 challenge on Text to Image (T2I) generation model quality assessment, which will be held in conjunction with the New Trends in Image Restoration and Enhancement Workshop (NTIRE) at CVPR 2025. The aim of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Shuhao Han , Haotian Fan , Fangyuan Kong , Wenjie Liao , Chunle Guo , Chongyi Li , Radu Timofte , Liang Li , Tao Li , Junhui Cui , Yunqiu Wang , Yang Tai , Jingwei Sun , Jianhui Sun , Xinli Yue , Tianyi Wang , Huan Hou , Junda Lu , Xinyang Huang , Zitang Zhou , Zijian Zhang , Xuhui Zheng , Xuecheng Wu , Chong Peng , Xuezhi Cao , Trong-Hieu Nguyen-Mau , Minh-Hoang Le , Minh-Khoa Le-Phan , Duy-Nam Ly , Hai-Dang Nguyen , Minh-Triet Tran , Yukang Lin , Yan Hong , Chuanbiao Song , Siyuan Li , Jun Lan , Zhichao Zhang , Xinyue Li , Wei Sun , Zicheng Zhang , Yunhao Li , Xiaohong Liu , Guangtao Zhai , Zitong Xu , Huiyu Duan , Jiarui Wang , Guangji Ma , Liu Yang , Lu Liu , Qiang Hu , Xiongkuo Min , Zichuan Wang , Zhenchen Tang , Bo Peng , Jing Dong , Fengbin Guan , Zihao Yu , Yiting Lu , Wei Luo , Xin Li , Minhao Lin , Haofeng Chen , Xuanxuan He , Kele Xu , Qisheng Xu , Zijian Gao , Tianjiao Wan , Bo-Cheng Qiu , Chih-Chung Hsu , Chia-ming Lee , Yu-Fan Lin , Bo Yu , Zehao Wang , Da Mu , Mingxiu Chen , Junkang Fang , Huamei Sun , Wending Zhao , Zhiyu Wang , Wang Liu , Weikang Yu , Puhong Duan , Bin Sun , Xudong Kang , Shutao Li , Shuai He , Lingzhi Fu , Heng Cong , Rongyu Zhang , Jiarong He , Zhishan Qiao , Yongqing Huang , Zewen Chen , Zhe Pang , Juan Wang , Jian Guo , Zhizhuo Shao , Ziyu Feng , Bing Li , Weiming Hu , Hesong Li , Dehua Liu , Zeming Liu , Qingsong Xie , Ruichen Wang , Zhihao Li , Yuqi Liang , Jianqi Bi , Jun Luo , Junfeng Yang , Can Li , Jing Fu , Hongwei Xu , Mingrui Long , Lulin Tang

How humans can effectively and efficiently acquire images has always been a perennial question. A classic solution is text-to-image retrieval from an existing database; however, the limited database typically lacks creativity. By contrast,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Leigang Qu , Haochuan Li , Tan Wang , Wenjie Wang , Yongqi Li , Liqiang Nie , Tat-Seng Chua

Dish images play a crucial role in the digital era, with the demand for culturally distinctive dish images continuously increasing due to the digitization of the food industry and e-commerce. In general cases, existing text-to-image…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Huijie Liu , Bingcan Wang , Jie Hu , Xiaoming Wei , Guoliang Kang

Food is essential to human survival. So much so that we have developed different recipes to suit our taste needs. In this work, we propose a novel way of creating new, fine-dining recipes from scratch using Transformers, specifically…

Computation and Language · Computer Science 2022-09-27 Konstantinos Katserelis , Konstantinos Skianis

This work presents an open-source unified benchmarking and evaluation framework for text-to-image generation models, with a particular focus on the impact of metadata augmented prompts. Leveraging the DeepFashion-MultiModal dataset, we…

Graphics · Computer Science 2025-05-09 Kapil Wanaskar , Gaytri Jena , Magdalini Eirinaki

Generating consistent long videos is a complex challenge: while diffusion-based generative models generate visually impressive short clips, extending them to longer durations often leads to memory bottlenecks and long-term inconsistency. In…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Wenqi Ouyang , Zeqi Xiao , Danni Yang , Yifan Zhou , Shuai Yang , Lei Yang , Jianlou Si , Xingang Pan

We propose a computational approach for recipe ideation, a downstream task that helps users select and gather ingredients for creating dishes. To perform this task, we developed RecipeMind, a food affinity score prediction model that…

Information Retrieval · Computer Science 2022-10-20 Mogan Gim , Donghee Choi , Kana Maruyama , Jihun Choi , Hajung Kim , Donghyeon Park , Jaewoo Kang

Significant progress has been made in the field of Instruction-based Image Editing (IIE). However, evaluating these models poses a significant challenge. A crucial requirement in this field is the establishment of a comprehensive evaluation…

Computer Vision and Pattern Recognition · Computer Science 2024-09-30 Yiwei Ma , Jiayi Ji , Ke Ye , Weihuang Lin , Zhibin Wang , Yonghan Zheng , Qiang Zhou , Xiaoshuai Sun , Rongrong Ji

The automatic recognition of food on images has numerous interesting applications, including nutritional tracking in medical cohorts. The problem has received significant research attention, but an ongoing public benchmark to develop open…

Identity-preserving text-to-video (IPT2V) generation creates videos faithful to both a reference subject image and a text prompt. While fine-tuning large pretrained video diffusion models on ID-matched data achieves state-of-the-art results…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jiayi Gao , Changcheng Hua , Qingchao Chen , Yuxin Peng , Yang Liu

Text-image-to-video (TI2V) generation is a critical problem for controllable video generation using both semantic and visual conditions. Most existing methods typically add visual conditions to text-to-video (T2V) foundation models by…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Bolin Lai , Sangmin Lee , Xu Cao , Xiang Li , James M. Rehg

Recipe is a set of instructions that describes how to make food. It can help people from the preparation of ingredients, food cooking process, etc. to prepare the food, and increasingly in demand on the Web. To help users find the vast…

Multimedia · Computer Science 2023-10-25 Jialiang Shi , Takahiro Komamizu , Keisuke Doman , Haruya Kyutoku , Ichiro Ide

State-of-the-art T2I models are capable of generating high-resolution images given textual prompts. However, they still struggle with accurately depicting compositional scenes that specify multiple objects, attributes, and spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Yixin Wan , Kai-Wei Chang

We present a new multimodal dataset called Visual Recipe Flow, which enables us to learn each cooking action result in a recipe text. The dataset consists of object state changes and the workflow of the recipe text. The state change is…

Computation and Language · Computer Science 2022-09-14 Keisuke Shirai , Atsushi Hashimoto , Taichi Nishimura , Hirotaka Kameko , Shuhei Kurita , Yoshitaka Ushiku , Shinsuke Mori

Text-to-video diffusion models enable the generation of high-quality videos that follow text instructions, making it easy to create diverse and individual content. However, existing approaches mostly focus on high-quality short video…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Roberto Henschel , Levon Khachatryan , Hayk Poghosyan , Daniil Hayrapetyan , Vahram Tadevosyan , Zhangyang Wang , Shant Navasardyan , Humphrey Shi