English
Related papers

Related papers: Analyzing Risky Choices: Q-Learning for Deal-No De…

200 papers

Across many domains of interaction, both natural and artificial, individuals use past experience to shape future behaviors. The results of such learning processes depend on what individuals wish to maximize. A natural objective is one's own…

Populations and Evolution · Quantitative Biology 2022-09-02 Alex McAvoy , Julian Kates-Harbeck , Krishnendu Chatterjee , Christian Hilbe

Consider a two-player zero-sum stochastic game where the transition function can be embedded in a given feature space. We propose a two-player Q-learning algorithm for approximating the Nash equilibrium strategy via sampling. The algorithm…

Machine Learning · Computer Science 2019-06-04 Zeyu Jia , Lin F. Yang , Mengdi Wang

In this paper, we explore the susceptibility of the independent Q-learning algorithms (a classical and widely used multi-agent reinforcement learning method) to strategic manipulation of sophisticated opponents in normal-form games played…

Computer Science and Game Theory · Computer Science 2024-07-17 Yuksel Arslantas , Ege Yuceel , Muhammed O. Sayin

We study sequential social learning with endogenous information acquisition when agents have a taste for nonconformity. Each agent observes predecessors' actions, chooses whether to acquire a private signal (and its precision), and then…

Theoretical Economics · Economics 2026-01-05 Georgy Lukyanov , Vasilii Ivanik

Q-learning is a popular Reinforcement Learning (RL) algorithm which is widely used in practice with function approximation (Mnih et al., 2015). In contrast, existing theoretical results are pessimistic about Q-learning. For example, (Baird,…

Machine Learning · Computer Science 2021-10-20 Naman Agarwal , Syomantak Chaudhuri , Prateek Jain , Dheeraj Nagaraj , Praneeth Netrapalli

Much research has been done to analyze the stock market. After all, if one can determine a pattern in the chaotic frenzy of transactions, then they could make a hefty profit from capitalizing on these insights. As such, the goal of our…

Machine Learning · Computer Science 2025-05-27 Ziyi Zhou , Nicholas Stern , Julien Laasri

In clinical practice, physicians make a series of treatment decisions over the course of a patient's disease based on his/her baseline and evolving characteristics. A dynamic treatment regime is a set of sequential decision rules that…

Methodology · Statistics 2015-02-04 Phillip J. Schulte , Anastasios A. Tsiatis , Eric B. Laber , Marie Davidian

The primary goal of reinforcement learning is to develop decision-making policies that prioritize optimal performance without considering risk or safety. In contrast, safe reinforcement learning aims to mitigate or avoid unsafe states. This…

Machine Learning · Computer Science 2024-09-13 Zahra Shahrooei , Ali Baheri

We develop methodology for a multistage decision problem with flexible number of stages in which the rewards are survival times that are subject to censoring. We present a novel Q-learning algorithm that is adjusted for censored data and…

Statistics Theory · Mathematics 2012-05-31 Yair Goldberg , Michael R. Kosorok

Banks routinely use neural networks to make decisions. While these models offer higher accuracy, they are susceptible to adversarial attacks, a risk often overlooked in the context of event sequences, particularly sequences of financial…

Following the risk-taking model of Seel and Strack, $n$ players decide when to stop privately observed Brownian motions with drift and absorption at zero. They are then ranked according to their level of stopping and paid a rank-dependent…

Optimization and Control · Mathematics 2021-11-09 Marcel Nutz , Yuchong Zhang

The article describes the use of deep Q-learning models in the problems of sales time series analytics. In contrast to supervised machine learning which is a kind of passive learning using historical data, Q-learning is a kind of active…

Machine Learning · Computer Science 2022-01-07 Bohdan M. Pavlyshenko

We study a model of irreversible investment for a decision-maker who has the possibility to gradually invest in a project with unknown value. In this setting, we introduce and explore a feature of "learning-by-doing", where the learning…

Optimization and Control · Mathematics 2024-06-25 Erik Ekström , Yerkin Kitapbayev , Alessandro Milazzo , Topias Tolonen-Weckström

This paper examines the convergence of no-regret learning in games with continuous action sets. For concreteness, we focus on learning via "dual averaging", a widely used class of no-regret learning schemes where players take small steps…

Optimization and Control · Mathematics 2018-01-17 Panayotis Mertikopoulos , Zhengyuan Zhou

In response to a change, individuals may choose to follow the responses of their friends or, alternatively, to change their friends. To model these decisions, consider a game where players choose their behaviors and friendships. In…

Social and Information Networks · Computer Science 2020-06-23 Anton Badev

Risk-averse total-reward Markov Decision Processes (MDPs) offer a promising framework for modeling and solving undiscounted infinite-horizon objectives. Existing model-based algorithms for risk measures like the entropic risk measure (ERM)…

Machine Learning · Computer Science 2025-10-27 Xihong Su , Jia Lin Hau , Gersi Doko , Kishan Panaganti , Marek Petrik

We consider long-lived agents who interact repeatedly in a social network. In each period, each agent learns about an unknown state by observing a private signal and her neighbors' actions from the previous period before choosing her own…

Theoretical Economics · Economics 2025-08-19 Florian Brandl

We consider risk-averse learning in repeated unknown games where the goal of the agents is to minimize their individual risk of incurring significantly high cost. Specifically, the agents use the conditional value at risk (CVaR) as a risk…

Machine Learning · Computer Science 2022-09-08 Zifan Wang , Yi Shen , Zachary I. Bell , Scott Nivison , Michael M. Zavlanos , Karl H. Johansson

We consider an online stochastic game with risk-averse agents whose goal is to learn optimal decisions that minimize the risk of incurring significantly high costs. Specifically, we use the Conditional Value at Risk (CVaR) as a risk measure…

Machine Learning · Computer Science 2022-06-17 Zifan Wang , Yi Shen , Michael M. Zavlanos

In many real world applications, reinforcement learning agents have to optimize multiple objectives while following certain rules or satisfying a list of constraints. Classical methods based on reward shaping, i.e. a weighted combination of…

Machine Learning · Computer Science 2020-09-15 Gabriel Kalweit , Maria Huegle , Moritz Werling , Joschka Boedecker