English

Verifier-Guided Twelve-Tone Composition: A Generate-Verify-Repair Harness for Symbolic Music Generation

Artificial Intelligence 2026-07-13 v1

Abstract

Large language models can produce superficially legal twelve-tone scores that collapse into degenerate textures. We introduce a neuro-symbolic harness that wraps a language-model proposer in a generate-verify-repair-trace loop with symbolic verification. The complete pipeline improves event-local consistency without claiming whole-piece legality. Across 40 controlled tasks and four paired models, audited delivery yield rises from 13.3% under raw generation to 48.1% with the harness, which explicitly abstains otherwise. The pass rate of a narrower collision and serialisation-consistency check rises from 33.5% to 58.3%, while degeneracy remains near 0.05, including under exploratory adversarial prompting. A blinded evaluation by five experts also shows a descriptive aggregate preference for harness candidates over raw generation in adherence, perceived legality, coherence, and overall quality.

Keywords

Cite

@article{arxiv.2607.11334,
  title  = {Verifier-Guided Twelve-Tone Composition: A Generate-Verify-Repair Harness for Symbolic Music Generation},
  author = {Congren Dai and Danni Zhao and Enyang Liu and Michael Ching Yam and Zhancheng Guo and Siyi Gu and Wentao Yang and Bo Dai and Xiaobing Li and Maosong Sun},
  journal= {arXiv preprint arXiv:2607.11334},
  year   = {2026}
}