中文
相关论文

相关论文: Structured Best Arm Identification with Fixed Conf…

200 篇论文

Imagine we want to split a group of agents into teams in the most \emph{efficient} way, considering that each agent has their own preferences about their teammates. This scenario is modeled by the extensively studied \textsc{Coalition…

数据结构与算法 · 计算机科学 2025-05-29 Foivos Fioravantes , Harmender Gahlawat , Nikolaos Melissinos

In multi-objective decision-making with hierarchical preferences, lexicographic bandits provide a natural framework for optimizing multiple objectives in a prioritized order. In this setting, a learner repeatedly selects arms and observes…

机器学习 · 计算机科学 2025-11-11 Bo Xue , Yuanyu Wan , Zhichao Lu , Qingfu Zhang

In this work, we address the problem of precisely localizing key frames of an action, for example, the precise time that a pitcher releases a baseball, or the precise time that a crowd begins to applaud. Key frame localization is a largely…

计算机视觉与模式识别 · 计算机科学 2020-01-22 Iljung S. Kwak , Jian-Zhong Guo , Adam Hantman , David Kriegman , Kristin Branson

We consider a finite-armed structured bandit problem in which mean rewards of different arms are known functions of a common hidden parameter $\theta^*$. Since we do not place any restrictions of these functions, the problem setting…

机器学习 · 统计学 2021-02-04 Samarth Gupta , Shreyas Chaudhari , Subhojyoti Mukherjee , Gauri Joshi , Osman Yağan

Neural networks are hypothesized to implement interpretable causal mechanisms, yet verifying this requires finding a causal abstraction -- a simpler, high-level Structural Causal Model (SCM) faithful to the network under interventions.…

机器学习 · 计算机科学 2026-03-02 Amir Asiaee

This paper proposes a new framework for providing approximation guarantees of local search algorithms. Local search is a basic algorithm design technique and is widely used for various combinatorial optimization problems. To analyze local…

数据结构与算法 · 计算机科学 2020-06-03 Kaito Fujii

We introduce the safe best-arm identification framework with linear feedback, where the agent is subject to some stage-wise safety constraint that linearly depends on an unknown parameter vector. The agent must take actions in a…

机器学习 · 统计学 2023-09-19 Xuedong Shang , Igor Colin , Merwan Barlier , Hamza Cherkaoui

As the adoption of federated learning increases for learning from sensitive data local to user devices, it is natural to ask if the learning can be done using implicit signals generated as users interact with the applications of interest,…

机器学习 · 计算机科学 2023-03-21 Alekh Agarwal , H. Brendan McMahan , Zheng Xu

We develop a mechanistic dynamical-systems formulation of best response in finite-action games with relational structure on the action set. The proposed neuromorphic decision dynamics realize best response as the stable outcome of an…

动力系统 · 数学 2026-05-05 Himani Sinhmar , Vaibhav Srivastava , Naomi Ehrich Leonard

We study the problem of estimating a continuous ability parameter from sequential binary responses by actively asking questions with varying difficulties, a setting that arises naturally in adaptive testing and online preference learning.…

机器学习 · 统计学 2025-10-10 Sanghwa Kim , Dohyun Ahn , Seungki Min

This paper introduces alignment games, a new class of zero-sum games modeling strategic interventions where effectiveness depends on alignment with an underlying hidden state. Motivated by operational problems in medical diagnostics,…

最优化与控制 · 数学 2025-09-08 Pedro Cesar Lopes Gerum , Thomas Lidbetter

We present an algorithm, "constrained successive accept or reject (CSAR)," for the problem of identifying the subset of top feasible-arms from a given finite set of arms with the limited sampling-budget equal to a given time-horizon when…

最优化与控制 · 数学 2025-01-22 Hyeong Soo Chang

We investigate the increasingly important and common game-solving setting where we do not have an explicit description of the game but only oracle access to it through gameplay, such as in financial or military simulations and computer…

人工智能 · 计算机科学 2020-02-26 Carlos Martin , Tuomas Sandholm

We study the problem of off-policy evaluation in the multi-armed bandit model with bounded rewards, and develop minimax rate-optimal procedures under three settings. First, when the behavior policy is known, we show that the Switch…

机器学习 · 统计学 2021-01-20 Cong Ma , Banghua Zhu , Jiantao Jiao , Martin J. Wainwright

We consider how an agent should update her beliefs when her beliefs are represented by a set P of probability distributions, given that the agent makes decisions using the minimax criterion, perhaps the best-studied and most commonly-used…

人工智能 · 计算机科学 2014-01-17 Peter D Grunwald , Joseph Y Halpern

We consider the problem of multi-fidelity zeroth-order optimization, where one can evaluate a function $f$ at various approximation levels (of varying costs), and the goal is to optimize $f$ with the cheapest evaluations possible. In this…

机器学习 · 计算机科学 2024-10-14 Étienne de Montbrun , Sébastien Gerchinovitz

We formulate, analyze and solve the problem of best arm identification with fairness constraints on subpopulations (BAICS). Standard best arm identification problems aim at selecting an arm that has the largest expected reward where the…

机器学习 · 计算机科学 2023-04-11 Yuhang Wu , Zeyu Zheng , Tingyu Zhu

We study fixed-confidence best-arm identification (BAI) where a cheap but potentially biased proxy (e.g., LLM judge) is available for every sample, while an expensive ground-truth label can only be acquired selectively when using a human…

机器学习 · 计算机科学 2026-01-30 Ruicheng Ao , Hongyu Chen , Siyang Gao , Hanwei Li , David Simchi-Levi

What can an agent learn in a stochastic Multi-Armed Bandit (MAB) problem from a dataset that contains just a single sample for each arm? Surprisingly, in this work, we demonstrate that even in such a data-starved setting it may still be…

机器学习 · 计算机科学 2024-02-27 Ruiqi Zhang , Yuexiang Zhai , Andrea Zanette

We study the representative arm identification (RAI) problem in the multi-armed bandits (MAB) framework, wherein we have a collection of arms, each associated with an unknown reward distribution. An underlying instance is defined by a…

机器学习 · 计算机科学 2024-08-27 Sarvesh Gharat , Aniket Yadav , Nikhil Karamchandani , Jayakrishnan Nair