English

Statistical Consequences of Dueling Bandits

Machine Learning 2021-11-02 v1 Statistics Theory Statistics Theory

Abstract

Multi-Armed-Bandit frameworks have often been used by researchers to assess educational interventions, however, recent work has shown that it is more beneficial for a student to provide qualitative feedback through preference elicitation between different alternatives, making a dueling bandits framework more appropriate. In this paper, we explore the statistical quality of data under this framework by comparing traditional uniform sampling to a dueling bandit algorithm and find that dueling bandit algorithms perform well at cumulative regret minimisation, but lead to inflated Type-I error rates and reduced power under certain circumstances. Through these results we provide insight into the challenges and opportunities in using dueling bandit algorithms to run adaptive experiments.

Keywords

Cite

@article{arxiv.2111.00870,
  title  = {Statistical Consequences of Dueling Bandits},
  author = {Nayan Saxena and Pan Chen and Emmy Liu},
  journal= {arXiv preprint arXiv:2111.00870},
  year   = {2021}
}

Comments

In Workshop on Reinforcement Learning for Education, 14th International Conference on Educational Data Mining , Paris, France, 2021

R2 v1 2026-06-24T07:20:45.199Z