中文
相关论文

相关论文: A Recipe for Creating Multimodal Aligned Datasets …

200 篇论文

The objective of this paper is a temporal alignment network that ingests long term video sequences, and associated text sentences, in order to: (1) determine if a sentence is alignable with the video; and (2) if it is alignable, then…

计算机视觉与模式识别 · 计算机科学 2022-04-07 Tengda Han , Weidi Xie , Andrew Zisserman

Many text generation applications require the generated text to be factually consistent with input information. Automatic evaluation of factual consistency is challenging. Previous work has developed various metrics that often depend on…

计算与语言 · 计算机科学 2023-05-29 Yuheng Zha , Yichi Yang , Ruichen Li , Zhiting Hu

Procedural video understanding is gaining attention in the vision and language community. Deep learning-based video analysis requires extensive data. Consequently, existing works often use web videos as training resources, making it…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Koki Maeda , Tosho Hirasawa , Atsushi Hashimoto , Jun Harashima , Leszek Rybicki , Yusuke Fukasawa , Yoshitaka Ushiku

This paper studies the challenge of developing robots capable of understanding under-specified instructions for creating functional object arrangements, such as "set up a dining table for two"; previous arrangement approaches have focused…

机器人学 · 计算机科学 2025-05-12 Yiqing Xu , Jiayuan Mao , Yilun Du , Tomas Lozáno-Pérez , Leslie Pack Kaelbling , David Hsu

We propose a computational approach for recipe ideation, a downstream task that helps users select and gather ingredients for creating dishes. To perform this task, we developed RecipeMind, a food affinity score prediction model that…

信息检索 · 计算机科学 2022-10-20 Mogan Gim , Donghee Choi , Kana Maruyama , Jihun Choi , Hajung Kim , Donghyeon Park , Jaewoo Kang

Human communication takes many forms, including speech, text and instructional videos. It typically has an underlying structure, with a starting point, ending, and certain objective steps between them. In this paper, we consider…

计算机视觉与模式识别 · 计算机科学 2016-05-12 Ozan Sener , Amir Roshan Zamir , Chenxia Wu , Silvio Savarese , Ashutosh Saxena

Cross-Modal sponsored search displays multi-modal advertisements (ads) when consumers look for desired products by natural language queries in search engines. Since multi-modal ads bring complementary details for query-ads matching, the…

计算机视觉与模式识别 · 计算机科学 2023-09-29 Yuanmin Tang , Jing Yu , Keke Gai , Yujing Wang , Yue Hu , Gang Xiong , Qi Wu

As a vast number of ingredients exist in the culinary world, there are countless food ingredient pairings, but only a small number of pairings have been adopted by chefs and studied by food researchers. In this work, we propose KitcheNette…

机器学习 · 计算机科学 2019-08-20 Donghyeon Park , Keonwoo Kim , Yonggyu Park , Jungwoon Shin , Jaewoo Kang

Recipe generation from food images and ingredients is a challenging task, which requires the interpretation of the information from another modality. Different from the image captioning task, where the captions usually have one sentence,…

计算机视觉与模式识别 · 计算机科学 2022-02-17 Hao Wang , Guosheng Lin , Steven C. H. Hoi , Chunyan Miao

In many applications involving multi-media data, the definition of similarity between items is integral to several key tasks, e.g., nearest-neighbor retrieval, classification, and recommendation. Data in such regimes typically exhibits…

人工智能 · 计算机科学 2010-09-01 Brian McFee , Gert Lanckriet

Comparing a user video to a reference how-to video is a key requirement for AR/VR technology delivering personalized assistance tailored to the user's progress. However, current approaches for language-based assistance can only answer…

计算机视觉与模式识别 · 计算机科学 2024-07-01 Tushar Nagarajan , Lorenzo Torresani

Existing large language models (LLMs) for machine translation are typically fine-tuned on sentence-level translation instructions and achieve satisfactory performance at the sentence level. However, when applied to document-level…

计算与语言 · 计算机科学 2024-01-17 Yachao Li , Junhui Li , Jing Jiang , Min Zhang

Learning to localize temporal boundaries of procedure steps in instructional videos is challenging due to the limited availability of annotated large-scale training videos. Recent works focus on learning the cross-modal alignment between…

计算机视觉与模式识别 · 计算机科学 2024-09-25 Yuxiao Chen , Kai Li , Wentao Bao , Deep Patel , Yu Kong , Martin Renqiang Min , Dimitris N. Metaxas

Creating an image reflecting the content of a long text is a complex process that requires a sense of creativity. For example, creating a book cover or a movie poster based on their summary or a food image based on its recipe. In this paper…

计算机视觉与模式识别 · 计算机科学 2019-01-09 Ori Bar El , Ori Licht , Netanel Yosephian

Tables stored in databases and tables which are present in web pages and articles account for a large part of semi-structured data that is available on the internet. It then becomes pertinent to develop a modeling approach with large…

计算与语言 · 计算机科学 2023-10-03 Soumajyoti Sarkar , Leonard Lausen

To acquire instruction-following capabilities, large language models (LLMs) undergo instruction tuning, where they are trained on instruction-response pairs using next-token prediction (NTP). Efforts to improve instruction tuning often…

计算与语言 · 计算机科学 2026-04-21 Yuxin Xiao , Shujian Zhang , Wenxuan Zhou , Marzyeh Ghassemi , Sanqiang Zhao

Due to limited supervised training data, large language models (LLMs) are typically pre-trained via a self-supervised "predict the next word" objective on a vast amount of unstructured text data. To make the resulting model useful to users,…

计算与语言 · 计算机科学 2026-01-30 Ajay Patel , Colin Raffel , Chris Callison-Burch

Procedural videos, exemplified by recipe demonstrations, are instrumental in conveying step-by-step instructions. However, understanding such videos is challenging as it involves the precise localization of steps and the generation of…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Anil Batra , Davide Moltisanti , Laura Sevilla-Lara , Marcus Rohrbach , Frank Keller

Recipe personalization through ingredient substitution has the potential to help people meet their dietary needs and preferences, avoid potential allergens, and ease culinary exploration in everyone's kitchen. To address ingredient…

机器学习 · 计算机科学 2023-02-17 Bahare Fatemi , Quentin Duval , Rohit Girdhar , Michal Drozdzal , Adriana Romero-Soriano

Procedural texts help AI enhance reasoning about context and action sequences. Transforming these into Semantic Role Labeling (SRL) improves understanding of individual steps by identifying predicate-argument structure like…

计算与语言 · 计算机科学 2025-05-28 Anil Batra , Laura Sevilla-Lara , Marcus Rohrbach , Frank Keller