Statistical Consequences of Dueling Bandits
Abstract
Multi-Armed-Bandit frameworks have often been used by researchers to assess educational interventions, however, recent work has shown that it is more beneficial for a student to provide qualitative feedback through preference elicitation between different alternatives, making a dueling bandits framework more appropriate. In this paper, we explore the statistical quality of data under this framework by comparing traditional uniform sampling to a dueling bandit algorithm and find that dueling bandit algorithms perform well at cumulative regret minimisation, but lead to inflated Type-I error rates and reduced power under certain circumstances. Through these results we provide insight into the challenges and opportunities in using dueling bandit algorithms to run adaptive experiments.
Cite
@article{arxiv.2111.00870,
title = {Statistical Consequences of Dueling Bandits},
author = {Nayan Saxena and Pan Chen and Emmy Liu},
journal= {arXiv preprint arXiv:2111.00870},
year = {2021}
}
Comments
In Workshop on Reinforcement Learning for Education, 14th International Conference on Educational Data Mining , Paris, France, 2021