中文
相关论文

相关论文: Near-Optimal Blacklisting

200 篇论文

Designing robust reinforcement learning (RL) agents in the presence of imperfect reward signals remains a core challenge. In practice, agents are often trained with proxy rewards that only approximate the true objective, leaving them…

机器学习 · 计算机科学 2026-04-15 Zixuan Liu , Xiaolin Sun , Zizhan Zheng

Black-box optimization is often encountered for decision-making in complex systems management, where the knowledge of system is limited. Under these circumstances, it is essential to balance the utilization of new information with…

统计计算 · 统计学 2025-01-15 Teng Lian , Jian-Qiang Hu , Yuhang Wu , Zeyu Zheng

The hidden-action model captures a fundamental problem of principal-agent theory and provides an optimal sharing rule when only the outcome but not the effort can be observed. However, the hidden-action model builds on various explicit and…

综合经济学 · 经济学 2020-04-15 Stephan Leitner , Friederike Wall

Current reinforcement learning methods fail if the reward function is imperfect, i.e. if the agent observes reward different from what it actually receives. We study this problem within the formalism of Corrupt Reward Markov Decision…

机器学习 · 计算机科学 2019-07-02 Jason Mancuso , Tomasz Kisielewski , David Lindner , Alok Singh

This paper considers a half-duplex scenario where an interferer behaves according to a parametric model but the values of the model parameters are unknown. We explore the necessary number of sensing steps to gather sufficient knowledge…

信息论 · 计算机科学 2024-10-11 Vincent Corlay , Jean-Christophe Sibel , Nicolas Gresset

Vacant taxi drivers' passenger seeking process in a road network generates additional vehicle miles traveled, adding congestion and pollution into the road network and the environment. This paper aims to employ a Markov Decision Process…

机器学习 · 计算机科学 2020-02-04 Zhenyu Shou , Xuan Di , Jieping Ye , Hongtu Zhu , Hua Zhang , Robert Hampshire

We study the iterative refinement of path planning for multiple robots, known as multi-agent pathfinding (MAPF). Given a graph, agents, their initial locations, and destinations, a solution of MAPF is a set of paths without collisions.…

机器人学 · 计算机科学 2022-02-15 Keisuke Okumura , Yasumasa Tamura , Xavier Defago

We consider the sequential decision-making problem of making proactive request assignment and rejection decisions for a profit-maximizing operator of an autonomous mobility on demand system. We formalize this problem as a Markov decision…

机器学习 · 计算机科学 2023-05-11 Tobias Enders , James Harrison , Marco Pavone , Maximilian Schiffer

Many real-world applications, such as those in medical domains, recommendation systems, etc, can be formulated as large state space reinforcement learning problems with only a small budget of the number of policy changes, i.e., low…

机器学习 · 计算机科学 2021-01-05 Minbo Gao , Tianle Xie , Simon S. Du , Lin F. Yang

A law in a multiagent system is a set of constraints imposed on agents' behaviours to avoid undesirable outcomes. The paper considers two types of laws: useful laws that, if followed, completely eliminate the undesirable outcomes and…

多智能体系统 · 计算机科学 2026-01-13 Qi Shi , Pavel Naumov

We consider the problem of selecting a subset of alternatives given noisy evaluations of the relative strength of different alternatives. We wish to select a k-subset (for a given k) that provides a maximum likelihood estimate for one of…

人工智能 · 计算机科学 2012-10-19 Ariel D. Procaccia , Sashank J. Reddi , Nisarg Shah

In this work we study the optimal execution problem with multiplicative price impact in algorithm trading, when an agent holds an initial position of shares of a financial asset. The inter-selling-decision times are modelled by the arrival…

数理金融 · 定量金融 2018-05-04 Daniel Hernández-Hernández , Harold A. Moreno-Franco , José Luis Pérez

We study sequential decision-making when the agent's internal model class is misspecified. Within the infinite-horizon Berk-Nash framework, stable behavior arises as a fixed point: the agent acts optimally relative to a subjective model,…

计算机科学与博弈论 · 计算机科学 2026-03-17 Quanyan Zhu , Zhengye Han

Currently the Dempster-Shafer based algorithm and Uniform Random Probability based algorithm are the preferred method of resolving security games, in which defenders are able to identify attackers and only strategy remained ambiguous.…

人工智能 · 计算机科学 2015-08-11 Hossein Khani , Mohsen Afsharchi

In this paper, we consider reinforcement learning of Markov Decision Processes (MDP) with peak constraints, where an agent chooses a policy to optimize an objective and at the same time satisfy additional constraints. The agent has to take…

最优化与控制 · 数学 2019-12-09 Ather Gattami

A large number of optimization algorithms have been developed by researchers to solve a variety of complex problems in operations management area. We present a novel optimization algorithm belonging to the class of swarm intelligence…

适应与自组织系统 · 物理学 2016-08-05 Ilario De Vincenzo , Ilaria Giannoccaro , Giuseppe Carbone

We consider a utility maximization problem over partially observable Markov ON/OFF channels. In this network instantaneous channel states are never known, and at most one user is selected for service in every slot according to the partial…

最优化与控制 · 数学 2010-08-23 Chih-ping Li , Michael J. Neely

Recent works have shown that agents facing independent instances of a stochastic $K$-armed bandit can collaborate to decrease regret. However, these works assume that each agent always recommends their individual best-arm estimates to other…

机器学习 · 计算机科学 2022-03-02 Daniel Vial , Sanjay Shakkottai , R. Srikant

Offline reinforcement learning (RL) aims at learning an optimal strategy using a pre-collected dataset without further interactions with the environment. While various algorithms have been proposed for offline RL in the previous literature,…

机器学习 · 计算机科学 2023-03-02 Wei Xiong , Han Zhong , Chengshuai Shi , Cong Shen , Liwei Wang , Tong Zhang

We consider an online strategic classification problem where each arriving agent can manipulate their true feature vector to obtain a positive predicted label, while incurring a cost that depends on the amount of manipulation. The learner…

机器学习 · 计算机科学 2024-03-28 Lingqing Shen , Nam Ho-Nguyen , Khanh-Hung Giang-Tran , Fatma Kılınç-Karzan
‹ 上一页 1 8 9 10 下一页 ›