中文

Kvasir-VQA:面向胃肠道解剖的文本-图像对数据集

计算机视觉与模式识别 2024-11-01 v1 人工智能

摘要

我们介绍 Kvasir-VQA,这是一个源自 HyperKvasir 和 Kvasir-Instrument 数据集的扩展数据集,增添了问答标注以促进胃肠道(Gastrointestinal, GI)诊断中高级机器学习任务。该数据集包含 6,500 幅带标注的图像,涵盖 various GI tract 状况和外科器械,支持多种问答类型,包括是/否、选择、位置和数值计数。该数据集旨在应用于图像描述、视觉问答(Visual Question Answering, VQA)、基于文本的合成医学图像生成、目标检测和分类。我们的实验表明,该数据集在训练三项选定任务的模型方面效果显著,展示了在医学图像分析和诊断中的重要应用。我们还为每个任务呈现评估指标,凸显了数据集的可用性和多样性。该数据集及其支持性资源均可在 https://datasets.simula.no/kvasir-vqa 获取。

关键词

引用

@article{arxiv.2409.01437,
  title  = {Kvasir-VQA: A Text-Image Pair GI Tract Dataset},
  author = {Sushant Gautam and Andrea Storås and Cise Midoglu and Steven A. Hicks and Vajira Thambawita and Pål Halvorsen and Michael A. Riegler},
  journal= {arXiv preprint arXiv:2409.01437},
  year   = {2024}
}

备注

to be published in VLM4Bio 2024, part of the ACM Multimedia (ACM MM) conference 2024