English

An Effective Data Creation Pipeline to Generate High-quality Financial Instruction Data for Large Language Model

Computation and Language 2023-08-04 v1 Artificial Intelligence Machine Learning

Abstract

At the beginning era of large language model, it is quite critical to generate a high-quality financial dataset to fine-tune a large language model for financial related tasks. Thus, this paper presents a carefully designed data creation pipeline for this purpose. Particularly, we initiate a dialogue between an AI investor and financial expert using ChatGPT and incorporate the feedback of human financial experts, leading to the refinement of the dataset. This pipeline yielded a robust instruction tuning dataset comprised of 103k multi-turn chats. Extensive experiments have been conducted on this dataset to evaluate the model's performance by adopting an external GPT-4 as the judge. The promising experimental results verify that our approach led to significant advancements in generating accurate, relevant, and financial-style responses from AI models, and thus providing a powerful tool for applications within the financial sector.

Keywords

Cite

@article{arxiv.2308.01415,
  title  = {An Effective Data Creation Pipeline to Generate High-quality Financial Instruction Data for Large Language Model},
  author = {Ziao Wang and Jianning Wang and Junda Wu and Xiaofeng Zhang},
  journal= {arXiv preprint arXiv:2308.01415},
  year   = {2023}
}

Comments

(FinLLM 2023)@IJCAI 2023

R2 v1 2026-06-28T11:46:49.567Z