中文
相关论文

相关论文: DiNeR: a Large Realistic Dataset for Evaluating Co…

200 篇论文

In this study, we introduced a new benchmark consisting of a curated dataset and a defined evaluation process to assess the compositional reasoning capabilities of large language models within the chemistry domain. We designed and validated…

Artificial intelligence (AI) systems built on incomplete or biased data will often exhibit problematic outcomes. Current methods of data analysis, particularly before model development, are costly and not standardized. The Dataset Nutrition…

数据库 · 计算机科学 2018-05-11 Sarah Holland , Ahmed Hosny , Sarah Newman , Joshua Joseph , Kasia Chmielinski

Standard methods in deep learning for natural language processing fail to capture the compositional structure of human language that allows for systematic generalization outside of the training distribution. However, human learners readily…

机器学习 · 计算机科学 2019-05-27 Jake Russin , Jason Jo , Randall C. O'Reilly , Yoshua Bengio

We study the problem of dataset distillation - creating a small set of synthetic examples capable of training a good model. In particular, we study the problem of label distillation - creating synthetic labels for a small set of real…

机器学习 · 计算机科学 2020-12-15 Ondrej Bohdal , Yongxin Yang , Timothy Hospedales

Large Language Models (LLMs) depend on high-quality, domain-specific natural language datasets. This dependency is particularly pronounced in Requirements Engineering (RE), where core activities rely on textual artifacts such as…

软件工程 · 计算机科学 2026-04-23 Quim Motger , Carlota Catot , Xavier Franch

Current state-of-the-art image generation models such as Latent Diffusion Models (LDMs) have demonstrated the capacity to produce visually striking food-related images. However, these generated images often exhibit an artistic or surreal…

计算机视觉与模式识别 · 计算机科学 2023-12-07 Olivia Markham , Yuhao Chen , Chi-en Amy Tai , Alexander Wong

Large language models (LLMs) have demonstrated remarkable generalizability, such as understanding arbitrary entities and relations. Instruction tuning has proven effective for distilling LLMs into more cost-efficient models such as Alpaca…

计算与语言 · 计算机科学 2024-01-22 Wenxuan Zhou , Sheng Zhang , Yu Gu , Muhao Chen , Hoifung Poon

Compared with English, Chinese suffers from more grammatical ambiguities, like fuzzy word boundaries and polysemous words. In this case, contextual information is not sufficient to support Chinese named entity recognition (NER), especially…

计算与语言 · 计算机科学 2022-10-25 Qinghua Mao , Jiatong Li , Kui Meng

Dataset distillation has emerged as a strategy to overcome the hurdles associated with large datasets by learning a compact set of synthetic data that retains essential information from the original dataset. While distilled data can be used…

机器学习 · 计算机科学 2024-07-23 William Yang , Ye Zhu , Zhiwei Deng , Olga Russakovsky

Compositional generalization is a key ability of humans that enables us to learn new concepts from only a handful examples. Neural machine learning models, including the now ubiquitous Transformers, struggle to generalize in this way, and…

机器学习 · 计算机科学 2024-01-19 Tim Klinger , Luke Liu , Soham Dan , Maxwell Crouse , Parikshit Ram , Alexander Gray

Natural language information needs over symbolic music scores rarely reduce to a single step lookup. Many queries require compositional Music Information Retrieval (MIR) that extracts multiple pieces of evidence from structured notation and…

机器学习 · 计算机科学 2026-03-02 Boyang Wang , Yash Vishe , Xin Xu , Zachary Novack , Xunyi Jiang , Julian McAuley , Junda Wu

Novelty modeling and detection is a core topic in Natural Language Processing (NLP), central to numerous tasks such as recommender systems and automatic summarization. It involves identifying pieces of text that deviate in some way from…

计算与语言 · 计算机科学 2025-05-14 Florian Carichon , Romain Rampa , Golnoosh Farnadi

With the booming of personalized recipe sharing networks (e.g., Yummly), a deluge of recipes from different cuisines could be obtained easily. In this paper, we aim to solve a problem which many home-cooks encounter when searching for…

信息检索 · 计算机科学 2020-03-17 Meng Chen , Xiaoyi Jia , Elizabeth Gorbonos , Chnh T. Hong , Xiaohui Yu , Yang Liu

Disease name normalization is an important task in the medical domain. It classifies disease names written in various formats into standardized names, serving as a fundamental component in smart healthcare systems for various…

计算与语言 · 计算机科学 2024-06-14 Wenqian Cui , Xiangling Fu , Shaohui Liu , Mingjun Gu , Xien Liu , Ji Wu , Irwin King

This work addresses the problem of learning sparse representations of tensor data using structured dictionary learning. It proposes learning a mixture of separable dictionaries to better capture the structure of tensor data by generalizing…

机器学习 · 计算机科学 2020-06-16 Mohsen Ghassemi , Zahra Shakeri , Anand D. Sarwate , Waheed U. Bajwa

Meal recommendation, as a typical health-related recommendation task, contains complex relationships between users, courses, and meals. Among them, meal-course affiliation associates user-meal and user-course interactions. However, an…

信息检索 · 计算机科学 2024-04-30 Ming Li , Lin Li , Xiaohui Tao , Jimmy Xiangji Huang

In this article, we evaluate four Large Language Models (LLMs) and their effectiveness at retrieving data within a specialized Retrieval-Augmented Generation (RAG) system, using a comprehensive food composition database. Our method is…

计算与语言 · 计算机科学 2026-03-12 Maks Požarnik Vavken , Matevž Ogrinc , Tome Eftimov , Barbara Koroušić Seljak

As people become more aware of their food choices, food computation models have become increasingly popular in assisting people in maintaining healthy eating habits. For example, food recommendation systems analyze recipe instructions to…

计算与语言 · 计算机科学 2023-06-06 Revathy Venkataramanan , Kaushik Roy , Kanak Raj , Renjith Prasad , Yuxin Zi , Vignesh Narayanan , Amit Sheth

We equip a smaller Language Model to generalise to answering challenging compositional questions that have not been seen in training. To do so we propose a combination of multitask supervised pretraining on up to 93 tasks designed to…

计算与语言 · 计算机科学 2023-08-22 Tim Hartill , Neset Tan , Michael Witbrock , Patricia J. Riddle

Current research in food analysis primarily concentrates on tasks such as food recognition, recipe retrieval and nutrition estimation from a single image. Nevertheless, there is a significant gap in exploring the impact of food intake on…

多媒体 · 计算机科学 2024-09-26 Yinxuan Gui , Bin Zhu , Jingjing Chen , Chong-Wah Ngo , Yu-Gang Jiang