English

Mind the Quote: Enabling Quotation-Aware Dialogue in LLMs via Plug-and-Play Modules

Artificial Intelligence 2025-06-02 v1 Computation and Language

Abstract

Human-AI conversation frequently relies on quoting earlier text-"check it with the formula I just highlighted"-yet today's large language models (LLMs) lack an explicit mechanism for locating and exploiting such spans. We formalise the challenge as span-conditioned generation, decomposing each turn into the dialogue history, a set of token-offset quotation spans, and an intent utterance. Building on this abstraction, we introduce a quotation-centric data pipeline that automatically synthesises task-specific dialogues, verifies answer correctness through multi-stage consistency checks, and yields both a heterogeneous training corpus and the first benchmark covering five representative scenarios. To meet the benchmark's zero-overhead and parameter-efficiency requirements, we propose QuAda, a lightweight training-based method that attaches two bottleneck projections to every attention head, dynamically amplifying or suppressing attention to quoted spans at inference time while leaving the prompt unchanged and updating < 2.8% of backbone weights. Experiments across models show that QuAda is suitable for all scenarios and generalises to unseen topics, offering an effective, plug-and-play solution for quotation-aware dialogue.

Keywords

Cite

@article{arxiv.2505.24292,
  title  = {Mind the Quote: Enabling Quotation-Aware Dialogue in LLMs via Plug-and-Play Modules},
  author = {Yueqi Zhang and Peiwen Yuan and Shaoxiong Feng and Yiwei Li and Xinglin Wang and Jiayi Shi and Chuyi Tan and Boyuan Pan and Yao Hu and Kan Li},
  journal= {arXiv preprint arXiv:2505.24292},
  year   = {2025}
}
R2 v1 2026-07-01T02:50:02.189Z