中文
相关论文

相关论文: CookAnything: A Framework for Flexible and Consist…

200 篇论文

Recipe generation from food images and ingredients is a challenging task, which requires the interpretation of the information from another modality. Different from the image captioning task, where the captions usually have one sentence,…

计算机视觉与模式识别 · 计算机科学 2022-02-17 Hao Wang , Guosheng Lin , Steven C. H. Hoi , Chunyan Miao

Nowadays, driven by the increasing concern on diet and health, food computing has attracted enormous attention from both industry and research community. One of the most popular research topics in this domain is Food Retrieval, due to its…

计算机视觉与模式识别 · 计算机科学 2020-04-03 Han Fu , Rui Wu , Chenghao Liu , Jianling Sun

While modern diffusion models excel at generating high-quality and diverse images, they still struggle with high-fidelity compositional and multimodal control, particularly when users simultaneously specify text prompts, subject references,…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Yusuf Dalva , Guocheng Gordon Qian , Maya Goldenberg , Tsai-Shien Chen , Kfir Aberman , Sergey Tulyakov , Pinar Yanardag , Kuan-Chieh Jackson Wang

People enjoy food photography because they appreciate food. Behind each meal there is a story described in a complex recipe and, unfortunately, by simply looking at a food image we do not have access to its preparation process. Therefore,…

计算机视觉与模式识别 · 计算机科学 2019-06-18 Amaia Salvador , Michal Drozdzal , Xavier Giro-i-Nieto , Adriana Romero

Creating recipe images is a key challenge in food computing, with applications in culinary education and multimodal recipe assistants. However, existing datasets lack fine-grained alignment between recipe goals, step-wise instructions, and…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Ruoxuan Zhang , Jidong Gao , Bin Wen , Hongxia Xie , Chenming Zhang , Hong-Han Shuai , Wen-Huang Cheng

Real-world meal images often contain multiple food items, making reliable compositional food image generation important for applications such as image-based dietary assessment, where multi-food data augmentation is needed, and recipe…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Xinyue Pan , Yuhao Chen , Fengqing Zhu

In this work we propose a new computational framework, based on generative deep models, for synthesis of photo-realistic food meal images from textual descriptions of its ingredients. Previous works on synthesis of images from text…

计算机视觉与模式识别 · 计算机科学 2019-05-31 Fangda Han , Ricardo Guerrero , Vladimir Pavlovic

Food is significant to human daily life. In this paper, we are interested in learning structural representations for lengthy recipes, that can benefit the recipe generation and food cross-modal retrieval tasks. Different from the common…

计算机视觉与模式识别 · 计算机科学 2022-08-02 Hao Wang , Guosheng Lin , Steven C. H. Hoi , Chunyan Miao

We introduce region-specific image refinement as a dedicated problem setting: given an input image and a user-specified region (e.g., a scribble mask or a bounding box), the goal is to restore fine-grained details while keeping all…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Dewei Zhou , You Li , Zongxin Yang , Yi Yang

Diffusion models have revolutionized image generation and editing, producing state-of-the-art results in conditioned and unconditioned image synthesis. While current techniques enable user control over the degree of change in an image edit,…

计算机视觉与模式识别 · 计算机科学 2024-03-01 Eran Levin , Ohad Fried

Controllable image synthesis models allow creation of diverse images based on text instructions or guidance from a reference image. Recently, denoising diffusion probabilistic models have been shown to generate more realistic imagery than…

计算机视觉与模式识别 · 计算机科学 2022-12-06 Xihui Liu , Dong Huk Park , Samaneh Azadi , Gong Zhang , Arman Chopikyan , Yuxiao Hu , Humphrey Shi , Anna Rohrbach , Trevor Darrell

We present a method to generate a video sequence given a single image. Because items in an image can be animated in arbitrarily many different ways, we introduce as control signal a sequence of motion strokes. Such control signal can be…

图像与视频处理 · 电气工程与系统科学 2020-08-17 Qiyang Hu , Adrian Wälchli , Tiziano Portenier , Matthias Zwicker , Paolo Favaro

We present a novel method for aligning a sequence of instructions to a video of someone carrying out a task. In particular, we focus on the cooking domain, where the instructions correspond to the recipe. Our technique relies on an HMM to…

计算与语言 · 计算机科学 2015-03-16 Jonathan Malmaud , Jonathan Huang , Vivek Rathod , Nick Johnston , Andrew Rabinovich , Kevin Murphy

Building on the success of diffusion models, significant advancements have been made in multimodal image generation tasks. Among these, human image generation has emerged as a promising technique, offering the potential to revolutionize the…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Shiyue Zhang , Zheng Chong , Xi Lu , Wenqing Zhang , Haoxiang Li , Xujie Zhang , Jiehui Huang , Xiao Dong , Xiaodan Liang

Diffusion models have recently been employed to generate high-quality images, reducing the need for manual data collection and improving model generalization in tasks such as object detection, instance segmentation, and image perception.…

计算机视觉与模式识别 · 计算机科学 2024-12-03 You Li , Fan Ma , Yi Yang

The effective communication of procedural knowledge remains a significant challenge in natural language processing (NLP), as purely textual instructions often fail to convey complex physical actions and spatial relationships. We address…

计算与语言 · 计算机科学 2025-05-23 Jing Bi , Pinxin Liu , Ali Vosoughi , Jiarui Wu , Jinxi He , Chenliang Xu

This paper introduces CookingSense, a descriptive collection of knowledge assertions in the culinary domain extracted from various sources, including web data, scientific papers, and recipes, from which knowledge covering a broad range of…

人工智能 · 计算机科学 2024-08-13 Donghee Choi , Mogan Gim , Donghyeon Park , Mujeen Sung , Hyunjae Kim , Jaewoo Kang , Jihun Choi

Pose-guided video generation refers to controlling the motion of subjects in generated video through a sequence of poses. It enables precise control over subject motion and has important applications in animation. However, current…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Ruiyan Wang , Teng Hu , Kaihui Huang , Zihan Su , Ran Yi , Lizhuang Ma

Modern deep learning techniques have enabled advances in image-based dietary assessment such as food recognition and food portion size estimation. Valuable information on the types of foods and the amount consumed are crucial for prevention…

计算机视觉与模式识别 · 计算机科学 2021-02-02 Jiangpeng He , Runyu Mao , Zeman Shao , Janine L. Wright , Deborah A. Kerr , Carol J. Boushey , Fengqing Zhu

Recent advances in text-to-3D scene generation have demonstrated significant potential to transform content creation across multiple industries. Although the research community has made impressive progress in addressing the challenges of…