QuArch: A Question-Answering Dataset for AI Agents in Computer Architecture
Hardware Architecture
2025-01-07 v2 Artificial Intelligence
Machine Learning
Abstract
We introduce QuArch, a dataset of 1500 human-validated question-answer pairs designed to evaluate and enhance language models' understanding of computer architecture. The dataset covers areas including processor design, memory systems, and performance optimization. Our analysis highlights a significant performance gap: the best closed-source model achieves 84% accuracy, while the top small open-source model reaches 72%. We observe notable struggles in memory systems, interconnection networks, and benchmarking. Fine-tuning with QuArch improves small model accuracy by up to 8%, establishing a foundation for advancing AI-driven computer architecture research. The dataset and leaderboard are at https://harvard-edge.github.io/QuArch/.
Cite
@article{arxiv.2501.01892,
title = {QuArch: A Question-Answering Dataset for AI Agents in Computer Architecture},
author = {Shvetank Prakash and Andrew Cheng and Jason Yik and Arya Tschand and Radhika Ghosal and Ikechukwu Uchendu and Jessica Quaye and Jeffrey Ma and Shreyas Grampurohit and Sofia Giannuzzi and Arnav Balyan and Fin Amin and Aadya Pipersenia and Yash Choudhary and Ankita Nayak and Amir Yazdanbakhsh and Vijay Janapa Reddi},
journal= {arXiv preprint arXiv:2501.01892},
year = {2025}
}