中文
相关论文

相关论文: Explore-then-Commit Algorithms for Decentralized T…

200 篇论文

The Exploration-Exploitation tradeoff arises in Reinforcement Learning when one cannot tell if a policy is optimal. Then, there is a constant need to explore new actions instead of exploiting past experience. In practice, it is common to…

机器学习 · 计算机科学 2019-09-10 Lior Shani , Yonathan Efroni , Shie Mannor

We consider (random) strategic interactions in a large population consisting of a variety of players. A rational player chooses actions that maximize certain utility functions, while a behavioral player chooses actions based on preferences…

最优化与控制 · 数学 2026-02-16 Raghupati Vyas , Kousik Das , Veeraruna Kavitha , Souvik Roy

This paper seeks to establish a framework for directing a society of simple, specialized, self-interested agents to solve what traditionally are posed as monolithic single-agent sequential decision problems. What makes it challenging to use…

机器学习 · 计算机科学 2020-08-17 Michael Chang , Sidhant Kaushik , S. Matthew Weinberg , Thomas L. Griffiths , Sergey Levine

We investigate the problem of learning an equilibrium in a generalized two-sided matching market, where agents can adaptively choose their actions based on their assigned matches. Specifically, we consider a setting in which matched agents…

机器学习 · 计算机科学 2025-06-05 Andreas Athanasopoulos , Christos Dimitrakakis

This document analyzes price discovery in cryptocurrency markets by comparing centralized and decentralized exchanges, as well as spot and futures markets. The study focuses first on Ethereum (ETH) and then applies a similar approach to…

交易与市场微观结构 · 定量金融 2025-06-11 Juan Plazuelo Pascual , Carlos Tardon Rubio , Juan Toro Cebada , Angel Hernando Veciana

Advancements in digitization have enabled two sided manufacturing-as-a-service (MaaS) marketplaces which has significantly reduced product development time for designers. These platforms provide designers with access to manufacturing…

人工智能 · 计算机科学 2025-06-17 Deepak Pahwa

We study reinforcement learning from human feedback in general Markov decision processes, where agents learn from trajectory-level preference comparisons. A central challenge in this setting is to design algorithms that select informative…

机器学习 · 计算机科学 2025-12-05 Andreas Schlaginhaufen , Reda Ouhamma , Maryam Kamgarpour

Two issues of algorithmic collusion are addressed in this paper. First, we show that in a general class of symmetric games, including Prisoner's Dilemma, Bertrand competition, and any (nonlinear) mixture of first and second price auction,…

理论经济学 · 经济学 2024-09-05 Zhang Xu , Wei Zhao

Matching plays a vital role in the rational allocation of resources in many areas, ranging from market operation to people's daily lives. In economics, the term matching theory is coined for pairing two agents in a specific market to reach…

社会与信息网络 · 计算机科学 2021-03-17 Jing Ren , Feng Xia , Xiangtai Chen , Jiaying Liu , Mingliang Hou , Ahsan Shehzad , Nargiz Sultanova , Xiangjie Kong

In many two-sided markets, the parties to be matched have incomplete information about their characteristics. We consider the settings where the parties engaged are extremely patient and are interested in long-term partnerships. Hence, once…

计算机科学与博弈论 · 计算机科学 2019-08-30 Kartik Ahuja , Mihaela van der Schaar

The matching literature often recommends market centralization under the assumption that agents know their own preferences and that their preferences are fixed. We find counterevidence to this assumption in a quasi-experiment. In Germany's…

综合经济学 · 经济学 2022-06-07 Julien Grenet , YingHua He , Dorothea Kübler

A menu description exposes strategyproofness by presenting a mechanism to player $i$ in two steps. Step (1) uses others' reports to describe $i$'s menu of potential outcomes. Step (2) uses $i$'s report to select $i$'s favorite outcome from…

理论经济学 · 经济学 2025-10-10 Yannai A. Gonczarowski , Ori Heffetz , Clayton Thomas

The paper addresses the Multiplayer Multi-Armed Bandit (MMAB) problem, where $M$ decision makers or players collaborate to maximize their cumulative reward. When several players select the same arm, a collision occurs and no reward is…

机器学习 · 计算机科学 2019-10-29 Alexandre Proutiere , Po-An Wang

Multi-access edge computing (MEC) is a promising architecture to provide low-latency applications for future Internet of Things (IoT)-based network systems. Together with the increasing scholarly attention on task offloading, the problem of…

分布式、并行与集群计算 · 计算机科学 2020-08-27 Zheng Xiao , Dan He , Yu Chen , Anthony Theodore Chronopoulos , Schahram Dustdar , Jiayi Du

Reinforcement learning in partially observed Markov decision processes (POMDPs) faces two challenges. (i) It often takes the full history to predict the future, which induces a sample complexity that scales exponentially with the horizon.…

机器学习 · 计算机科学 2024-04-02 Lingxiao Wang , Qi Cai , Zhuoran Yang , Zhaoran Wang

Remote entanglement enables coordinated decision making without communication and produces correlations beyond those achievable by any classical strategy, representing a practical quantum advantage in time-critical distributed…

量子物理 · 物理学 2026-04-10 Changhao Li , Seigo Kikura , Akihisa Goban , Hayata Yamasaki , Shinichi Sunami

The threat of algorithmic collusion, and whether it merits regulatory intervention, remains debated, as existing evaluations of its emergence often rely on long learning horizons, assumptions about counterparty rationality in adopting…

多智能体系统 · 计算机科学 2026-03-11 Yuhong Luo , Daniel Schoepflin , Xintong Wang

We study an online mixed discrete and continuous optimization problem where a decision maker interacts with an unknown environment for a number of $T$ rounds. At each round, the decision maker needs to first jointly choose a discrete and a…

最优化与控制 · 数学 2024-08-27 Lintao Ye , Ming Chi , Zhi-Wei Liu , Xiaoling Wang , Vijay Gupta

One-sided matching problems with ordinal preferences, such as hostel room allocation, are commonly solved using the Top Trading Cycles (TTC) mechanism, which guarantees Pareto-optimal (PO) outcomes. However, TTC does not yield a unique…

计算机科学与博弈论 · 计算机科学 2026-05-14 Bhavik Dodda , Garima Shakya

Partial monitoring games are repeated games where the learner receives feedback that might be different from adversary's move or even the reward gained by the learner. Recently, a general model of combinatorial partial monitoring (CPM)…

计算机科学与博弈论 · 计算机科学 2016-08-24 Sougata Chaudhuri , Ambuj Tewari