English

COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning

Computation and Language 2024-11-05 v2 Artificial Intelligence

Abstract

Remarkable progress on English instruction tuning has facilitated the efficacy and reliability of large language models (LLMs). However, there remains a noticeable gap in instruction tuning for Chinese, where the complex linguistic features pose significant challenges. Existing datasets, generally distilled from English-centric LLMs, are not well-aligned with Chinese users' interaction patterns. To bridge this gap, we introduce COIG-CQIA, a new Chinese instruction tuning dataset derived from various real-world resources and undergoing rigorous human verification. We conduct extensive experiments on COIG-CQIA, and compare them with strong baseline models and datasets. The experimental results show that models trained on COIG-CQIA achieve highly competitive performance in diverse benchmarks. Additionally, our findings offer several insights for designing effective Chinese instruction-tuning datasets and data-mixing strategies. Our dataset are available at https://huggingface.co/datasets/m-a-p/COIG-CQIA.

Keywords

Cite

@article{arxiv.2403.18058,
  title  = {COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning},
  author = {Yuelin Bai and Xinrun Du and Yiming Liang and Yonggang Jin and Junting Zhou and Ziqiang Liu and Feiteng Fang and Mingshan Chang and Tianyu Zheng and Xincheng Zhang and Nuo Ma and Zekun Wang and Ruibin Yuan and Haihong Wu and Hongquan Lin and Wenhao Huang and Jiajun Zhang and Chenghua Lin and Jie Fu and Min Yang and Shiwen Ni and Ge Zhang},
  journal= {arXiv preprint arXiv:2403.18058},
  year   = {2024}
}