English
Related papers

Related papers: FMiFood: Multi-modal Contrastive Learning for Food…

200 papers

Accurate nutrient estimation from unstructured recipe text is an important yet challenging problem in dietary monitoring, due to ambiguous ingredient terminology and highly variable quantity expressions. We systematically evaluate models…

Computation and Language · Computer Science 2026-05-14 Wei-Chun Chen , Yu-Xuan Chen , I-Fang Chung , Ying-Jia Lin

Assessment of dietary intake has primarily relied on self-report instruments, which are prone to measurement errors. Dietary assessment methods have increasingly incorporated technological advances particularly mobile, image based…

Computer Vision and Pattern Recognition · Computer Science 2022-09-13 Gautham Vinod , Zeman Shao , Fengqing Zhu

With the rise and development of computer vision and LLMs, intelligence is everywhere, especially for people and cars. However, for tremendous food attributes (such as origin, quantity, weight, quality, sweetness, etc.), existing research…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Zhenbo Xu , Jinghan Yang , Gong Huang , Jiqing Feng , Liu Liu , Ruihan Sun , Ajin Meng , Zhuo Zhang , Zhaofeng He

Until recently, the number of public real-world text images was insufficient for training scene text recognizers. Therefore, most modern training methods rely on synthetic data and operate in a fully supervised manner. Nevertheless, the…

Computer Vision and Pattern Recognition · Computer Science 2022-05-10 Aviad Aberdam , Roy Ganz , Shai Mazor , Ron Litman

The Image Difference Captioning (IDC) task aims to describe the visual differences between two similar images with natural language. The major challenges of this task lie in two aspects: 1) fine-grained visual differences that require…

Multimedia · Computer Science 2022-02-10 Linli Yao , Weiying Wang , Qin Jin

We perform a comprehensive benchmarking of contrastive frameworks for learning multimodal representations in the medical domain. Through this study, we aim to answer the following research questions: (i) How transferable are general-domain…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Shuvendu Roy , Yasaman Parhizkar , Franklin Ogidi , Vahid Reza Khazaie , Michael Colacci , Ali Etemad , Elham Dolatabadi , Arash Afkanpour

Food classification from images is a fine-grained classification problem. Manual curation of food images is cost, time and scalability prohibitive. On the other hand, web data is available freely but contains noise. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2017-12-27 Parneet Kaur , Karan Sikka , Ajay Divakaran

In recent years, several unsupervised, "contrastive" learning algorithms in vision have been shown to learn representations that perform remarkably well on transfer tasks. We show that this family of algorithms maximizes a lower bound on…

Machine Learning · Computer Science 2020-06-08 Mike Wu , Chengxu Zhuang , Milan Mosse , Daniel Yamins , Noah Goodman

In the recent past, complex deep neural networks have received huge interest in various document understanding tasks such as document image classification and document retrieval. As many document types have a distinct visual style, learning…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Souhail Bakkali , Ziheng Ming , Mickael Coustaty , Marçal Rusiñol

Multimodal search has revolutionized the fashion industry, providing a seamless and intuitive way for users to discover and explore fashion items. Based on their preferences, style, or specific attributes, users can search for products by…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Prithviraj Purushottam Naik , Rohit Agarwal

Food recognition has a wide range of applications, such as health-aware recommendation and self-service restaurants. Most previous methods of food recognition firstly locate informative regions in some weakly-supervised manners and then…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Yaohui Zhu , Linhu Liu , Jiang Tian

Content-based medical image retrieval is an important diagnostic tool that improves the explainability of computer-aided diagnosis systems and provides decision making support to healthcare professionals. Medical imaging data, such as…

Computer Vision and Pattern Recognition · Computer Science 2022-11-23 Yunyan Xing , Benjamin J. Meyer , Mehrtash Harandi , Tom Drummond , Zongyuan Ge

Measuring biodiversity is crucial for understanding ecosystem health. While prior works have developed machine learning models for taxonomic classification of photographic images and DNA separately, in this work, we introduce a multimodal…

Artificial Intelligence · Computer Science 2025-12-10 ZeMing Gong , Austin T. Wang , Xiaoliang Huo , Joakim Bruslund Haurum , Scott C. Lowe , Graham W. Taylor , Angel X. Chang

With the rapid development of society and continuous advances in science and technology, the food industry increasingly demands higher production quality and efficiency. Food image classification plays a vital role in enabling automated…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Xinle Gao , Linghui Ye , Zhiyong Xiao

Automatic detection of multimodal fake news has gained a widespread attention recently. Many existing approaches seek to fuse unimodal features to produce multimodal news representations. However, the potential of powerful cross-modal…

Machine Learning · Computer Science 2023-08-14 Longzheng Wang , Chuang Zhang , Hongbo Xu , Yongxiu Xu , Xiaohan Xu , Siqi Wang

Few-shot object detection (FSOD) has garnered significant research attention in the field of remote sensing due to its ability to reduce the dependency on large amounts of annotated data. However, two challenges persist in this area: (1)…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Jiawei Zhou , Wuzhou Li , Yi Cao , Hongtao Cai , Xiang Li

Accurate molecular property prediction requires integrating complementary information from molecular structure and chemical semantics. In this work, we propose LGM-CL, a local-global multimodal contrastive learning framework that jointly…

Machine Learning · Computer Science 2026-02-02 Xiayu Liu , Zhengyi Lu , Yunhong Liao , Chan Fan , Hou-biao Li

Contrastive language-image Pre-training (CLIP) [13] can leverage large datasets of unlabeled Image-Text pairs, which have demonstrated impressive performance in various downstream tasks. Given that annotating medical data is time-consuming…

Image and Video Processing · Electrical Eng. & Systems 2023-07-13 Yuhao Wang

Image-text retrieval is a central problem for understanding the semantic relationship between vision and language, and serves as the basis for various visual and language tasks. Most previous works either simply learn coarse-grained…

Computer Vision and Pattern Recognition · Computer Science 2023-07-19 Chong Liu , Yuqi Zhang , Hongsong Wang , Weihua Chen , Fan Wang , Yan Huang , Yi-Dong Shen , Liang Wang

Automated food intake gesture detection plays a vital role in dietary monitoring, enabling objective and continuous tracking of eating behaviors to support better health outcomes. Wrist-worn inertial measurement units (IMUs) have been…

Machine Learning · Computer Science 2025-07-11 Chunzhuo Wang , Hans Hallez , Bart Vanrumste