Graphusion: Leveraging Large Language Models for Scientific Knowledge Graph Fusion and Construction in NLP Education
Abstract
Knowledge graphs (KGs) are crucial in the field of artificial intelligence and are widely applied in downstream tasks, such as enhancing Question Answering (QA) systems. The construction of KGs typically requires significant effort from domain experts. Recently, Large Language Models (LLMs) have been used for knowledge graph construction (KGC), however, most existing approaches focus on a local perspective, extracting knowledge triplets from individual sentences or documents. In this work, we introduce Graphusion, a zero-shot KGC framework from free text. The core fusion module provides a global view of triplets, incorporating entity merging, conflict resolution, and novel triplet discovery. We showcase how Graphusion could be applied to the natural language processing (NLP) domain and validate it in the educational scenario. Specifically, we introduce TutorQA, a new expert-verified benchmark for graph reasoning and QA, comprising six tasks and a total of 1,200 QA pairs. Our evaluation demonstrates that Graphusion surpasses supervised baselines by up to 10% in accuracy on link prediction. Additionally, it achieves average scores of 2.92 and 2.37 out of 3 in human evaluations for concept entity extraction and relation recognition, respectively.
Keywords
Cite
@article{arxiv.2407.10794,
title = {Graphusion: Leveraging Large Language Models for Scientific Knowledge Graph Fusion and Construction in NLP Education},
author = {Rui Yang and Boming Yang and Sixun Ouyang and Tianwei She and Aosong Feng and Yuang Jiang and Freddy Lecue and Jinghui Lu and Irene Li},
journal= {arXiv preprint arXiv:2407.10794},
year = {2024}
}
Comments
24 pages, 11 figures, 13 tables. arXiv admin note: substantial text overlap with arXiv:2402.14293