中文
相关论文

相关论文: Multi-modal Cooking Workflow Construction for Food…

200 篇论文

Multimodal large language models (MLLMs) are rapidly evolving, presenting increasingly complex safety challenges. However, current dataset construction methods, which are risk-oriented, fail to cover the growing complexity of real-world…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Jingen Qu , Lijun Li , Bo Zhang , Yichen Yan , Jing Shao

Workflows provide an expressive programming model for fine-grained control of large-scale applications in distributed computing environments. Accurate estimates of complex workflow execution metrics on large-scale machines have several key…

分布式、并行与集群计算 · 计算机科学 2018-04-18 Alok Singh , Mai Nguyen , Shweta Purawat , Daniel Crawl , Ilkay Altintas

Recent advancements in Vision-Language Models (VLMs) have revolutionized general visual understanding. However, their application in the food domain remains constrained by benchmarks that rely on coarse-grained categories, single-view…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Song Jin , Juntian Zhang , Xun Zhang , Zeying Tian , Fei Jiang , Guojun Yin , Wei Lin , Yong Liu , Rui Yan

The reproducibility issue in science has come under increased scrutiny. One consistent suggestion lies in the use of scripted methods or workflows for data analysis. Image analysis is one area in science in which little can be done in…

图像与视频处理 · 电气工程与系统科学 2017-09-22 Paul A. Thompson , Norm Matloff , Alex Fu , Ariel Shin

Understanding behavior requires datasets that capture humans while carrying out complex tasks. The kitchen is an excellent environment for assessing human motor and cognitive function, as many complex actions are naturally exhibited in…

Interests in the automatic generation of cooking recipes have been growing steadily over the past few years thanks to a large amount of online cooking recipes. We present RecipeGPT, a novel online recipe generation and evaluation system.…

计算与语言 · 计算机科学 2020-10-28 Helena H. Lee , Ke Shu , Palakorn Achananuparp , Philips Kokoh Prasetyo , Yue Liu , Ee-Peng Lim , Lav R. Varshney

Multimodal large models have shown great potential in automating pathology image analysis. However, current multimodal models for gastrointestinal pathology are constrained by both data quality and reasoning transparency: pervasive noise…

图像与视频处理 · 电气工程与系统科学 2025-07-25 Minxi Ouyang , Lianghui Zhu , Yaqing Bao , Qiang Huang , Jingli Ouyang , Tian Guan , Xitong Ling , Jiawen Li , Song Duan , Wenbin Dai , Li Zheng , Xuemei Zhang , Yonghong He

We introduce the World Wide recipe, which sets forth a framework for culturally aware and participatory data collection, and the resultant regionally diverse World Wide Dishes evaluation dataset. We also analyse bias operationalisation to…

With the advent of the era of foundation models, pre-training and fine-tuning have become common paradigms. Recently, parameter-efficient fine-tuning has garnered widespread attention due to its better balance between the number of…

计算机视觉与模式识别 · 计算机科学 2024-08-02 Bin Cheng , Jiaxuan Lu

The reasoning segmentation task, which demands a nuanced comprehension of intricate queries to accurately pinpoint object regions, is attracting increasing attention. However, Multi-modal Large Language Models (MLLM) often find it difficult…

计算机视觉与模式识别 · 计算机科学 2024-07-12 Xiaoyi Bao , Siyang Sun , Shuailei Ma , Kecheng Zheng , Yuxin Guo , Guosheng Zhao , Yun Zheng , Xingang Wang

Node graph systems are used ubiquitously for material design in computer graphics. They allow the use of visual programming to achieve desired effects without writing code. As high-level design tools they provide convenience and…

图形学 · 计算机科学 2023-04-27 Yiwei Hu , Paul Guerrero , Miloš Hašan , Holly Rushmeier , Valentin Deschaintre

Food recognition is one of the most important components in image-based dietary assessment. However, due to the different complexity level of food images and inter-class similarity of food categories, it is challenging for an image-based…

计算机视觉与模式识别 · 计算机科学 2020-12-08 Runyu Mao , Jiangpeng He , Zeman Shao , Sri Kalyan Yarlagadda , Fengqing Zhu

Nutrition estimation of meals from visual data is an important problem for dietary monitoring and computational health, but existing approaches largely rely on single images of the finally completed dish. This setting is fundamentally…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Chengkun Yue , Chuanzhi Xu , Jiangpeng He

Conversational Recommender Systems (CRSs) aim to provide personalized recommendations by interacting with users through conversations. Most existing studies of CRS focus on extracting user preferences from conversational contexts. However,…

信息检索 · 计算机科学 2025-04-28 Yibiao Wei , Jie Zou , Weikang Guo , Guoqing Wang , Xing Xu , Yang Yang

People often create art by following an artistic workflow involving multiple stages that inform the overall design. If an artist wishes to modify an earlier decision, significant work may be required to propagate this new decision forward…

计算机视觉与模式识别 · 计算机科学 2020-07-15 Hung-Yu Tseng , Matthew Fisher , Jingwan Lu , Yijun Li , Vladimir Kim , Ming-Hsuan Yang

Multi-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance, a key capability for future AI assistants. Effective guidance requires not only delivering…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Apratim Bhattacharyya , Bicheng Xu , Sanjay Haresh , Reza Pourreza , Litian Liu , Sunny Panchal , Pulkit Madan , Leonid Sigal , Roland Memisevic

It will be increasingly common for robots to operate in cluttered human-centered environments such as homes, workplaces, and hospitals, where the robot is often tasked to maintain perception constraints, such as monitoring people or…

机器人学 · 计算机科学 2026-03-05 Qingxi Meng , Emiliano Flores , Thai Duong , Vaibhav Unhelkar , Lydia E. Kavraki

The goal of this work is to generate step-by-step visual instructions in the form of a sequence of images, given an input image that provides the scene context and the sequence of textual instructions. This is a challenging problem as it…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Tomáš Souček , Prajwal Gatti , Michael Wray , Ivan Laptev , Dima Damen , Josef Sivic

This paper is based on developing different algorithms, which generate the task tree planning for the given goal node(recipe). The knowledge representation of the dishes is called FOON. It contains the different objects and their between…

机器人学 · 计算机科学 2023-12-18 Chakradhar Reddy Nallu

We established a rigorous benchmark for text-based recipe generation, a fundamental task in natural language generation. We present a comprehensive comparative study contrasting a fine-tuned GPT-2 large (774M) model against the GPT-2 small…

计算与语言 · 计算机科学 2025-08-21 Shubham Pundhir , Ganesh Bagler