English
Related papers

Related papers: RecipeGen: A Step-Aligned Multimodal Benchmark for…

200 papers

Recipe image generation is an important challenge in food computing, with applications from culinary education to interactive recipe platforms. However, there is currently no real-world dataset that comprehensively connects recipe goals,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Ruoxuan Zhang , Hongxia Xie , Yi Yao , Jian-Yu Jiang-Lin , Bin Wen , Ling Lo , Hong-Han Shuai , Yung-Hui Li , Wen-Huang Cheng

We present Step-Video-TI2V, a state-of-the-art text-driven image-to-video generation model with 30B parameters, capable of generating videos up to 102 frames based on both text and image inputs. We build Step-Video-TI2V-Eval as a new…

The burgeoning field of Artificial Intelligence Generated Content (AIGC) is witnessing rapid advancements, particularly in video generation. This paper introduces AIGCBench, a pioneering comprehensive and scalable benchmark designed to…

Computer Vision and Pattern Recognition · Computer Science 2024-01-24 Fanda Fan , Chunjie Luo , Wanling Gao , Jianfeng Zhan

Significant work has been conducted in the domain of food computing, yet these studies typically focus on single tasks such as t2t (instruction generation from food titles and ingredients), i2t (recipe generation from food images), or t2i…

Computer Vision and Pattern Recognition · Computer Science 2024-09-19 Peiyu Li , Xiaobao Huang , Yijun Tian , Nitesh V. Chawla

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also neglect this…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Kaiyue Sun , Kaiyi Huang , Xian Liu , Yue Wu , Zihan Xu , Zhenguo Li , Xihui Liu

Recent advances in text-to-video (T2V) technology, as demonstrated by models such as Runway Gen-3, Pika, Sora, and Kling, have significantly broadened the applicability and popularity of the technology. This progress has created a growing…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Zelu Qi , Ping Shi , Shuqi Wang , Chaoyang Zhang , Fei Zhao , Zefeng Ying , Da Pan , Xi Yang , Zheqi He , Teng Dai

Text-to-video (T2V) models have shown remarkable performance in generating visually reasonable scenes, while their capability to leverage world knowledge for ensuring semantic consistency and factual accuracy remains largely understudied.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Yubin Chen , Xuyang Guo , Zhenmei Shi , Zhao Song , Jiahao Zhang

Querying generative AI models, e.g., large language models (LLMs), has become a prevalent method for information acquisition. However, existing query-answer datasets primarily focus on textual responses, making it challenging to address…

Artificial Intelligence · Computer Science 2025-06-03 Shuting Wang , Yunqi Liu , Zixin Yang , Ning Hu , Zhicheng Dou , Chenyan Xiong

Learning effective recipe representations is essential in food studies. Unlike what has been developed for image-based recipe retrieval or learning structural text embeddings, the combined effect of multi-modal information (i.e., recipe…

Machine Learning · Computer Science 2022-05-26 Yijun Tian , Chuxu Zhang , Zhichun Guo , Yihong Ma , Ronald Metoyer , Nitesh V. Chawla

Recent advances in text-to-image (T2I) generation have achieved impressive results, yet existing models still struggle with prompts that require rich world knowledge and implicit reasoning: both of which are critical for producing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Daoan Zhang , Che Jiang , Ruoshi Xu , Biaoxiang Chen , Zijian Jin , Yutian Lu , Jianguo Zhang , Liang Yong , Jiebo Luo , Shengda Luo

Food image segmentation is a critical and indispensible task for developing health-related applications such as estimating food calories and nutrients. Existing food image segmentation models are underperforming due to two reasons: (1)…

Computer Vision and Pattern Recognition · Computer Science 2021-05-13 Xiongwei Wu , Xin Fu , Ying Liu , Ee-Peng Lim , Steven C. H. Hoi , Qianru Sun

Video generation models are revolutionizing content creation, with image-to-video models drawing increasing attention due to their enhanced controllability, visual consistency, and practical applications. However, despite their popularity,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Wenhao Wang , Yi Yang

As large language models have demonstrated impressive performance in many domains, recent works have adopted language models (LMs) as controllers of visual modules for vision-and-language tasks. While existing work focuses on equipping LMs…

Computer Vision and Pattern Recognition · Computer Science 2023-10-30 Jaemin Cho , Abhay Zala , Mohit Bansal

Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and video in isolation or rely on coarse embedding similarity,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Ziwei Zhou , Zeyuan Lai , Rui Wang , Yifan Yang , Zhen Xing , Yuqing Yang , Qi Dai , Lili Qiu , Chong Luo

Thanks to recent advancements in scalable deep architectures and large-scale pretraining, text-to-video generation has achieved unprecedented capabilities in producing high-fidelity, instruction-following content across a wide range of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Xuyang Guo , Jiayan Huo , Zhenmei Shi , Zhao Song , Jiahao Zhang , Jiale Zhao

Recent advancements in text-to-image (T2I) generation have enabled models to produce high-quality images from textual descriptions. However, these models often struggle with complex instructions involving multiple objects, attributes, and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Yucheng Zhou , Jiahao Yuan , Qianning Wang

Advances in technology have led to the development of methods that can create desired visual multimedia. In particular, image generation using deep learning has been extensively studied across diverse fields. In comparison, video…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Doyeon Kim , Donggyu Joo , Junmo Kim

Generative models have demonstrated remarkable capability in synthesizing high-quality text, images, and videos. For video generation, contemporary text-to-video models exhibit impressive capabilities, crafting visually stunning videos.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Jay Zhangjie Wu , Guian Fang , Haoning Wu , Xintao Wang , Yixiao Ge , Xiaodong Cun , David Junhao Zhang , Jia-Wei Liu , Yuchao Gu , Rui Zhao , Weisi Lin , Wynne Hsu , Ying Shan , Mike Zheng Shou

Infographics are composite visual artifacts that combine data visualizations with textual and illustrative elements to communicate information. While recent text-to-image (T2I) models can generate aesthetically appealing images, their…

Despite the impressive advances in text-to-image models, they often struggle to effectively compose complex scenes with multiple objects, displaying various attributes and relationships. To address this challenge, we present…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Kaiyi Huang , Chengqi Duan , Kaiyue Sun , Enze Xie , Zhenguo Li , Xihui Liu
‹ Prev 1 2 3 10 Next ›