English

Enhancing Systematic Reviews with Large Language Models: Using GPT-4 and Kimi

Computation and Language 2025-04-30 v1 Applications

Abstract

This research delved into GPT-4 and Kimi, two Large Language Models (LLMs), for systematic reviews. We evaluated their performance by comparing LLM-generated codes with human-generated codes from a peer-reviewed systematic review on assessment. Our findings suggested that the performance of LLMs fluctuates by data volume and question complexity for systematic reviews.

Keywords

Cite

@article{arxiv.2504.20276,
  title  = {Enhancing Systematic Reviews with Large Language Models: Using GPT-4 and Kimi},
  author = {Dandan Chen Kaptur and Yue Huang and Xuejun Ryan Ji and Yanhui Guo and Bradley Kaptur},
  journal= {arXiv preprint arXiv:2504.20276},
  year   = {2025}
}

Comments

13 pages, Paper presented at the National Council on Measurement in Education (NCME) Conference, Denver, Colorado, in April 2025

R2 v1 2026-06-28T23:14:32.525Z