English

ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation

Machine Learning 2025-09-30 v1 Artificial Intelligence Computation and Language

Abstract

We introduce ORPO-Distill, a general-purpose method for cross-architecture LLM distillation that formulates the problem as a preference optimization task. Unlike standard CoT distillation, the approach transfers knowledge through diverse reasoning traces. It employs an Odds-Ratio Preference Optimization objective that contrasts teacher and student traces for more effective learning, and adopts a mixed-policy strategy for utilizing student-generated outputs, outperforming both off- and on-policy alternatives. Experiments on five datasets and multiple student models show consistent improvements over conventional black-box KD baselines.

Keywords

Cite

@article{arxiv.2509.25100,
  title  = {ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation},
  author = {Aasheesh Singh and Vishal Vaddina and Dagnachew Birru},
  journal= {arXiv preprint arXiv:2509.25100},
  year   = {2025}
}

Comments

Accepted at NeurIPS 2025, Efficient Reasoning Workshop

R2 v1 2026-07-01T06:05:16.800Z