中文
相关论文

相关论文: Revamping Cross-Modal Recipe Retrieval with Hierar…

200 篇论文

Despite the abundance of multi-modal data, such as image-text pairs, there has been little effort in understanding the individual entities and their different roles in the construction of these data instances. In this work, we endeavour to…

计算机视觉与模式识别 · 计算机科学 2021-02-05 Hai X. Pham , Ricardo Guerrero , Jiatong Li , Vladimir Pavlovic

Cross-modal image-recipe retrieval has gained significant attention in recent years. Most work focuses on improving cross-modal embeddings using unimodal encoders, that allow for efficient retrieval in large-scale databases, leaving aside…

计算机视觉与模式识别 · 计算机科学 2022-04-22 Mustafa Shukor , Guillaume Couairon , Asya Grechka , Matthieu Cord

Designing powerful tools that support cooking activities has rapidly gained popularity due to the massive amounts of available data, as well as recent advances in machine learning that are capable of analyzing them. In this paper, we…

计算与语言 · 计算机科学 2018-05-01 Micael Carvalho , Rémi Cadène , David Picard , Laure Soulier , Nicolas Thome , Matthieu Cord

Food is significant to human daily life. In this paper, we are interested in learning structural representations for lengthy recipes, that can benefit the recipe generation and food cross-modal retrieval tasks. Different from the common…

计算机视觉与模式识别 · 计算机科学 2022-08-02 Hao Wang , Guosheng Lin , Steven C. H. Hoi , Chunyan Miao

This paper addresses the challenges of learning representations for recipes and food images in the cross-modal retrieval problem. As the relationship between a recipe and its cooked dish is cause-and-effect, treating a recipe as a text…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Qing Wang , Chong-Wah Ngo , Ee-Peng Lim

Food computing is playing an increasingly important role in human daily life, and has found tremendous applications in guiding human behavior towards smart food consumption and healthy lifestyle. An important task under the food-computing…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Hao Wang , Doyen Sahoo , Chenghao Liu , Ee-peng Lim , Steven C. H. Hoi

Food image-to-recipe aims to learn an embedded space linking the rich semantics in recipes with the visual content in food image for cross-modal retrieval. The existing research works carry out the learning of such space by assuming that…

多媒体 · 计算机科学 2023-04-18 Bin Zhu , Chong-Wah Ngo , Jingjing Chen , Wing-Kwong Chan

Existing approaches for image-to-recipe retrieval have the implicit assumption that a food image can fully capture the details textually documented in its recipe. However, a food image only reflects the visual outcome of a cooked dish and…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Qing Wang , Chong-Wah Ngo , Yu Cao , Ee-Peng Lim

We propose a novel non-parametric method for cross-modal recipe retrieval which is applied on top of precomputed image and text embeddings. By combining our method with standard approaches for building image and text encoders, trained…

计算机视觉与模式识别 · 计算机科学 2021-07-14 Mikhail Fain , Niall Twomey , Andrey Ponikar , Ryan Fox , Danushka Bollegala

Cross-modal recipe retrieval has attracted research attention in recent years, thanks to the availability of large-scale paired data for training. Nevertheless, obtaining adequate recipe-image pairs covering the majority of cuisines for…

计算机视觉与模式识别 · 计算机科学 2022-05-10 Bin Zhu , Chong-Wah Ngo , Jingjing Chen , Wing-Kwong Chan

In this paper, we present a cross-modal recipe retrieval framework, Transformer-based Network for Large Batch Training (TNLBT), which is inspired by ACME~(Adversarial Cross-Modal Embedding) and H-T~(Hierarchical Transformer). TNLBT aims to…

计算机视觉与模式识别 · 计算机科学 2022-12-19 Jing Yang , Junwen Chen , Keiji Yanai

Cross-modal retrieval between food images and recipe texts is an important task with applications in nutritional management, dietary logging, and cooking assistance. Existing methods predominantly rely on dual-encoder architectures with…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Keisuke Gomi , Keiji Yanai

Food retrieval is an important task to perform analysis of food-related information, where we are interested in retrieving relevant information about the queried food item such as ingredients, cooking instructions, etc. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2021-09-22 Hao Wang , Doyen Sahoo , Chenghao Liu , Ke Shu , Palakorn Achananuparp , Ee-peng Lim , Steven C. H. Hoi

In this paper, we introduce Recipe1M+, a new large-scale, structured corpus of over one million cooking recipes and 13 million food images. As the largest publicly available collection of recipe data, Recipe1M+ affords the ability to train…

计算机视觉与模式识别 · 计算机科学 2019-07-11 Javier Marin , Aritro Biswas , Ferda Ofli , Nicholas Hynes , Amaia Salvador , Yusuf Aytar , Ingmar Weber , Antonio Torralba

Computational food analysis (CFA) naturally requires multi-modal evidence of a particular food, e.g., images, recipe text, etc. A key to making CFA possible is multi-modal shared representation learning, which aims to create a joint…

计算机视觉与模式识别 · 计算机科学 2021-10-01 Ricardo Guerrero , Hai Xuan Pham , Vladimir Pavlovic

Direct computer vision based-nutrient content estimation is a demanding task, due to deformation and occlusions of ingredients, as well as high intra-class and low inter-class variability between meal classes. In order to tackle these…

信息检索 · 计算机科学 2019-11-06 Matthias Fontanellaz , Stergios Christodoulidis , Stavroula Mougiakakou

People enjoy food photography because they appreciate food. Behind each meal there is a story described in a complex recipe and, unfortunately, by simply looking at a food image we do not have access to its preparation process. Therefore,…

计算机视觉与模式识别 · 计算机科学 2019-06-18 Amaia Salvador , Michal Drozdzal , Xavier Giro-i-Nieto , Adriana Romero

Food is essential to human survival. So much so that we have developed different recipes to suit our taste needs. In this work, we propose a novel way of creating new, fine-dining recipes from scratch using Transformers, specifically…

计算与语言 · 计算机科学 2022-09-27 Konstantinos Katserelis , Konstantinos Skianis

Vision-language retrieval is an important multi-modal learning topic, where the goal is to retrieve the most relevant visual candidate for a given text query. Recently, pre-trained models, e.g., CLIP, show great potential on retrieval…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Haojun Jiang , Jianke Zhang , Rui Huang , Chunjiang Ge , Zanlin Ni , Shiji Song , Gao Huang

Current state-of-the-art approaches to cross-modal retrieval process text and visual input jointly, relying on Transformer-based architectures with cross-attention mechanisms that attend over all words and objects in an image. While…

计算机视觉与模式识别 · 计算机科学 2022-02-22 Gregor Geigle , Jonas Pfeiffer , Nils Reimers , Ivan Vulić , Iryna Gurevych
‹ 上一页 1 2 3 10 下一页 ›