English

Don't Pay Attention, PLANT It: Pretraining Attention via Learning-to-Rank

Computation and Language 2025-12-29 v4 Machine Learning

Abstract

State-of-the-art Extreme Multi-Label Text Classification models rely on multi-label attention to focus on key tokens in input text, but learning good attention weights is challenging. We introduce PLANT - Pretrained and Leveraged Attention - a plug-and-play strategy for initializing attention. PLANT works by planting label-specific attention using a pretrained Learning-to-Rank model guided by mutual information gain. This architecture-agnostic approach integrates seamlessly with large language model backbones such as Mistral-7B, LLaMA3-8B, DeepSeek-V3, and Phi-3. PLANT outperforms state-of-the-art methods across tasks including ICD coding, legal topic classification, and content recommendation. Gains are especially pronounced in few-shot settings, with substantial improvements on rare labels. Ablation studies confirm that attention initialization is a key driver of these gains. For code and trained models, see https://github.com/debjyotiSRoy/xcube/tree/plant

Keywords

Cite

@article{arxiv.2410.23066,
  title  = {Don't Pay Attention, PLANT It: Pretraining Attention via Learning-to-Rank},
  author = {Debjyoti Saha Roy and Byron C. Wallace and Javed A. Aslam},
  journal= {arXiv preprint arXiv:2410.23066},
  year   = {2025}
}