将多源测试时自适应作为抽取式问答中的对决老虎机问题
计算与语言
2023-06-13 v1
摘要
本工作中,我们研究基于用户反馈的多源测试时模型自适应,其中建立 K 个 distinct 模型用于自适应。为实现高效自适应,我们将该问题建模为随机决策过程,旨在确定自适应后最优的已自适应模型。我们讨论了两种框架:多臂老虎机学习(multi-armed bandit learning)与多臂对决老虎机(multi-armed dueling bandits)。相较于多臂老虎机学习,对决框架允许 K 个模型间的两两协作,这由本工作提出的名为 Co-UCB 的新方法求解。在六个抽取式问答 (QA) 数据集上的实验表明,使用 Co-UCB 的对决框架对于我们研究的问题比其他强基线更有效。
引用
@article{arxiv.2306.06779,
title = {Multi-Source Test-Time Adaptation as Dueling Bandits for Extractive Question Answering},
author = {Hai Ye and Qizhe Xie and Hwee Tou Ng},
journal= {arXiv preprint arXiv:2306.06779},
year = {2023}
}
备注
Main conference of ACL 2023