中文
相关论文

相关论文: Gamifying optimization: a Wasserstein distance-bas…

200 篇论文

We present a novel $Q$-learning algorithm tailored to solve distributionally robust Markov decision problems where the corresponding ambiguity set of transition probabilities for the underlying Markov decision process is a Wasserstein ball…

机器学习 · 计算机科学 2024-06-21 Ariel Neufeld , Julian Sester

Policy optimization is a core component of reinforcement learning (RL), and most existing RL methods directly optimize parameters of a policy based on maximizing the expected total reward, or its surrogate. Though often achieving…

机器学习 · 计算机科学 2018-08-10 Ruiyi Zhang , Changyou Chen , Chunyuan Li , Lawrence Carin

As opposed to standard empirical risk minimization (ERM), distributionally robust optimization aims to minimize the worst-case risk over a larger ambiguity set containing the original empirical distribution of the training data. In this…

机器学习 · 计算机科学 2021-01-06 Jaeho Lee , Maxim Raginsky

This article extends the idea of solving parity games by strategy iteration to non-deterministic strategies: In a non-deterministic strategy a player restricts himself to some non-empty subset of possible actions at a given node, instead of…

计算机科学与博弈论 · 计算机科学 2012-03-20 Michael Luttenberger

Unlike traditional time series, the action sequences of human decision making usually involve many cognitive processes such as beliefs, desires, intentions, and theory of mind, i.e., what others are thinking. This makes predicting human…

机器学习 · 计算机科学 2022-06-07 Baihan Lin , Djallel Bouneffouf , Guillermo Cecchi

We study the trade-offs between strategyproofness and other desiderata, such as efficiency or fairness, that often arise in the design of random ordinal mechanisms. We use approximate strategyproofness to define manipulability, a measure to…

计算机科学与博弈论 · 计算机科学 2017-01-11 Timo Mennle , Sven Seuken

Recent real-time heuristic search algorithms have demonstrated outstanding performance in video-game pathfinding. However, their applications have been thus far limited to that domain. We proceed with the aim of facilitating wider…

人工智能 · 计算机科学 2013-08-16 Daniel Huntley , Vadim Bulitko

Reinforcement learning algorithms rely on carefully engineering environment rewards that are extrinsic to the agent. However, annotating each environment with hand-designed, dense rewards is not scalable, motivating the need for developing…

机器学习 · 计算机科学 2018-08-14 Yuri Burda , Harri Edwards , Deepak Pathak , Amos Storkey , Trevor Darrell , Alexei A. Efros

In recent years, Wasserstein Distributionally Robust Optimization (DRO) has garnered substantial interest for its efficacy in data-driven decision-making under distributional uncertainty. However, limited research has explored the…

机器学习 · 计算机科学 2025-10-01 Ahmad-Reza Ehyaei , Golnoosh Farnadi , Samira Samadi

This monograph develops a comprehensive statistical learning framework that is robust to (distributional) perturbations in the data using Distributionally Robust Optimization (DRO) under the Wasserstein metric. Beginning with fundamental…

机器学习 · 统计学 2021-08-23 Ruidi Chen , Ioannis Ch. Paschalidis

In many areas of data mining, data is collected from humans beings. In this contribution, we ask the question of how people actually respond to ordinal scales. The main problem observed is that users tend to be volatile in their choices,…

人机交互 · 计算机科学 2017-03-01 Kevin Jasberg , Sergej Sizov

We study the problem of resource provisioning under stringent reliability or service level requirements, which arise in applications such as power distribution, emergency response, cloud server provisioning, and regulatory risk management.…

最优化与控制 · 数学 2025-04-11 Anand Deo , Karthyek Murthy

Many modern machine learning applications, such as multi-task learning, require finding optimal model parameters to trade-off multiple objective functions that may conflict with each other. The notion of the Pareto set allows us to focus on…

最优化与控制 · 数学 2022-09-05 Mao Ye , Qiang Liu

Existing observational approaches for learning human preferences, such as inverse reinforcement learning, usually make strong assumptions about the observability of the human's environment. However, in reality, people make many important…

机器学习 · 统计学 2021-10-29 Cassidy Laidlaw , Stuart Russell

Wagering mechanisms are one-shot betting mechanisms that elicit agents' predictions of an event. For deterministic wagering mechanisms, an existing impossibility result has shown incompatibility of some desirable theoretical properties. In…

计算机科学与博弈论 · 计算机科学 2022-03-31 Yiling Chen , Yang Liu , Juntao Wang

Behavioral experiments on the Ultimatum Game have shown that we human beings have remarkable preference in fair play, contradicting the predictions by the game theory. Most of the existing models seeking for explanations, however, strictly…

种群与进化 · 定量生物学 2022-12-07 Guozhong Zheng , Jiqiang Zhang , Rizhou Liang , Lin Ma , Li Chen

Many studies have shown that there are regularities in the way human beings make decisions. However, our ability to obtain models that capture such regularities and can accurately predict unobserved decisions is still limited. We tackle…

综合金融 · 定量金融 2021-03-11 Gael Poux-Medard , Sergio Cobo-Lopez , Jordi Duch , Roger Guimera , Marta Sales-Pardo

Machine learning applications frequently come with multiple diverse objectives and constraints that can change over time. Accordingly, trained models can be tuned with sets of hyper-parameters that affect their predictive behavior (e.g.,…

机器学习 · 计算机科学 2022-10-17 Bracha Laufer-Goldshtein , Adam Fisch , Regina Barzilay , Tommi Jaakkola

Decision tree learning is a widely used approach in machine learning, favoured in applications that require concise and interpretable models. Heuristic methods are traditionally used to quickly produce models with reasonably high accuracy.…

Strategic decision-making in uncertain and adversarial environments is crucial for the security of modern systems and infrastructures. A salient feature of many optimal decision-making policies is a level of unpredictability, or randomness,…

计算机科学与博弈论 · 计算机科学 2024-05-03 Keith Paarporn , Rahul Chandan , Dan Kovenock , Mahnoosh Alizadeh , Jason R. Marden