引导隐藏状态:大型音频语言模型中链式思考推理的无训练模型操控
摘要
链式思考 (CoT) 提示已扩展到 large audio-language models (LALMs) 以激发 reasoning,但 enhance its effectiveness without training 仍具有挑战性。我们 study inference-time model steering 作为 training-free approach 来 improve LALM reasoning。我们引入 three strategies using diverse information sources 并 across four LALMs and four benchmarks 进行 evaluation。结果显示在 CoT prompting 下 general accuracy gains 最高可达 4.4%。 notably, we identify a cross-modal transfer where steering vectors derived from few text samples effectively guide speech-based reasoning, demonstrating high data efficiency。we also examine hyperparameter sensitivity to understand the robustness of these approaches。our findings position model steering as a practical direction for strengthening LALM reasoning。
引用
@article{arxiv.2603.14636,
title = {Nudging Hidden States: Training-Free Model Steering for Chain-of-Thought Reasoning in Large Audio-Language Models},
author = {Lok-Lam Ieong and Chia-Chien Chen and Chih-Kai Yang and Yu-Han Huang and An-Yu Cheng and Hung-yi Lee},
journal= {arXiv preprint arXiv:2603.14636},
year = {2026}
}
备注
6 pages, 4 figures, 2 tables