中文
相关论文

相关论文: UMDFood: Vision-language models boost food composi…

200 篇论文

Reliable Uncertainty Quantification (UQ) and failure prediction remain open challenges for Vision-Language Models (VLMs). We introduce ViLU, a new Vision-Language Uncertainty quantification framework that contextualizes uncertainty…

The success of large-scale visual language pretraining (VLP) models has driven widespread adoption of image-text retrieval tasks. However, their deployment on mobile devices remains limited due to large model sizes and computational…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Yuqi Li , Chuanguang Yang , Junhao Dong , Zhengtao Yao , Haoyan Xu , Zeyu Dong , Hansheng Zeng , Zhulin An , Yingli Tian

Existing Multimodal Large Language Models (MLLMs) follow the paradigm that perceives visual information by aligning visual features with the input space of Large Language Models (LLMs), and concatenating visual tokens with text tokens to…

计算机视觉与模式识别 · 计算机科学 2024-05-31 Feipeng Ma , Hongwei Xue , Guangting Wang , Yizhou Zhou , Fengyun Rao , Shilin Yan , Yueyi Zhang , Siying Wu , Mike Zheng Shou , Xiaoyan Sun

Online reviews in the form of user-generated content (UGC) significantly impact consumer decision-making. However, the pervasive issue of not only human fake content but also machine-generated content challenges UGC's reliability. Recent…

机器学习 · 计算机科学 2024-01-18 Alessandro Gambetti , Qiwei Han

Obesity treatment requires obese patients to record all food intakes per day. Computer vision has been introduced to estimate calories from food images. In order to increase accuracy of detection and reduce the error of volume estimation in…

计算机视觉与模式识别 · 计算机科学 2018-02-20 Yanchao Liang , Jianhua Li

Half of long-term care (LTC) residents are malnourished increasing hospitalization, mortality, morbidity, with lower quality of life. Current tracking methods are subjective and time consuming. This paper presents the automated food imaging…

计算机视觉与模式识别 · 计算机科学 2021-12-10 Kaylen J. Pfisterer , Robert Amelard , Jennifer Boger , Audrey G. Chung , Heather H. Keller , Alexander Wong

Image-to-recipe retrieval is a challenging vision-to-language task of significant practical value. The main challenge of the task lies in the ultra-high redundancy in the long recipe and the large variation reflected in both food item…

计算机视觉与模式识别 · 计算机科学 2023-05-22 Bhanu Prakash Voutharoja , Peng Wang , Lei Wang , Vivienne Guan

Automating the detection of fruits and vegetables using computer vision is essential for modernizing agriculture, improving efficiency, ensuring food quality, and contributing to technologically advanced and sustainable farming practices.…

计算机视觉与模式识别 · 计算机科学 2024-09-23 Sandeep Khanna , Chiranjoy Chattopadhyay , Suman Kundu

Food is essential for human survival, and people always try to taste different types of delicious recipes. Frequently, people choose food ingredients without even knowing their names or pick up some food ingredients that are not obvious to…

计算机视觉与模式识别 · 计算机科学 2022-03-15 Md. Shafaat Jamil Rokon , Md Kishor Morol , Ishra Binte Hasan , A. M. Saif , Rafid Hussain Khan

Food recognition is one of the most important components in image-based dietary assessment. However, due to the different complexity level of food images and inter-class similarity of food categories, it is challenging for an image-based…

计算机视觉与模式识别 · 计算机科学 2020-12-08 Runyu Mao , Jiangpeng He , Zeman Shao , Sri Kalyan Yarlagadda , Fengqing Zhu

We present a mobile application made to recognize food items of multi-object meal from a single image in real-time, and then return the nutrition facts with components and approximate amounts. Our work is organized in two parts. First, we…

计算机视觉与模式识别 · 计算机科学 2019-09-17 Jianing Sun , Katarzyna Radecka , Zeljko Zilic

Unified vision large language models (VLLMs) have recently achieved impressive advancements in both multimodal understanding and generation, powering applications such as visual question answering and text-guided image synthesis. However,…

计算与语言 · 计算机科学 2025-09-19 Pengyu Wang , Shaojun Zhou , Chenkun Tan , Xinghao Wang , Wei Huang , Zhen Ye , Zhaowei Li , Botian Jiang , Dong Zhang , Xipeng Qiu

Multi-view learning often faces challenges in effectively leveraging images captured from different angles and locations. This challenge is particularly pronounced when addressing inconsistencies and uncertainties between views. In this…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Jiwoong Yang , Haejun Chung , Ikbeom Jang

$ $As a result of bad eating habits, humanity may be destroyed. People are constantly on the lookout for tasty foods, with junk foods being the most common source. As a consequence, our eating patterns are shifting, and we're gravitating…

计算机视觉与模式识别 · 计算机科学 2022-03-23 Sirajum Munira Shifat , Takitazwar Parthib , Sabikunnahar Talukder Pyaasa , Nila Maitra Chaity , Niloy Kumar , Md. Kishor Morol

Detecting an ingestion environment is an important aspect of monitoring dietary intake. It provides insightful information for dietary assessment. However, it is a challenging problem where human-based reviewing can be tedious, and…

Accurate food nutrition estimation from single images is challenging due to the loss of 3D information. While depth-based methods provide reliable geometry, they remain inaccessible on most smartphones because of depth-sensor requirements.…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Darrin Bright , Rakshith Raj , Kanchan Keisham

Accurate food volume estimation is crucial for dietary monitoring, medical nutrition management, and food intake analysis. Existing 3D Food Volume estimation methods accurately compute the food volume but lack for food portions selection.…

图形学 · 计算机科学 2025-06-04 Ahmad AlMughrabi , Umair Haroon , Ricardo Marques , Petia Radeva

Vision-Language Models (VLMs) have demonstrated strong capabilities in aligning visual and textual modalities, enabling a wide range of applications in multimodal understanding and generation. While they excel in zero-shot and transfer…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Hao Dong , Moru Liu , Jian Liang , Eleni Chatzi , Olga Fink

Large Multimodal Models (LMMs) are increasingly applied to meal images for nutrition analysis. However, existing work primarily evaluates proprietary models, such as GPT-4. This leaves the broad range of LLMs underexplored. Additionally,…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Bruce Coburn , Jiangpeng He , Megan E. Rollo , Satvinder S. Dhaliwal , Deborah A. Kerr , Fengqing Zhu

Multimodal language generation, which leverages the synergy of language and vision, is a rapidly expanding field. However, existing vision-language models face challenges in tasks that require complex linguistic understanding. To address…

计算与语言 · 计算机科学 2023-12-20 Jiwan Chung , Youngjae Yu