肺癌知识库构建:探讨LLM中的语义推理
摘要
将大语言模型(LLM)集成到生物医学研究中为领域特定推理和知识表示带来了新机会。然而,其性能高度依赖训练数据的语义质量。在肿瘤学中,由于精确性和可解释性至关重要,构建结构化知识库的可扩展方法对于有效微调至关重要。本研究提出了开发肺癌知识库的管道,使用开放信息抽取(OpenIE)。该过程包括:(1)识别医学概念使用MeSH辞典;(2)筛选获批准许可的PubMed文献(CC0);(3)使用OpenIE方法提取(主语、谓语、宾语)三元组;(4)使用命名实体识别(NER)丰富三元组集合以确保生物医学相关性。 resulting triplet sets provide a domain-specific, large-scale, and noise-aware resource for fine-tuning LLMs。我们评估了在该数据集上进行监督式语义微调的T5模型。与ROUGE和BERTScore的比较评估显示,性能和语义连贯性显著提高,表明OpenIE派生资源作为提升生物医学NLP的可扩展、低成本解决方案具有潜力。
引用
@article{arxiv.2601.02604,
title = {Scalable Construction of a Lung Cancer Knowledge Base: Profiling Semantic Reasoning in LLMs},
author = {Cesar Felipe Martínez Cisneros and Jesús Ulises Quiroz Bautista and Claudia Anahí Guzmán Solano and Bogdan Kaleb García Rivera and Iván García Pacheco and Yalbi Itzel Balderas Martínez and Kolawole John Adebayoc and Ignacio Arroyo Fernández},
journal= {arXiv preprint arXiv:2601.02604},
year = {2026}
}
备注
\c{opyright} 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works