Related papers: New axioms for top trading cycles
This paper presents a new theory, known as robust dynamic pro- gramming, for a class of continuous-time dynamical systems. Different from traditional dynamic programming (DP) methods, this new theory serves as a fundamental tool to analyze…
Organisms have evolved a variety of mechanisms to cope with the unpredictability of environmental conditions, and yet mainstream models of metabolic regulation are typically based on strict optimality principles that do not account for…
Machine learning algorithms aim to find patterns from observations, which may include some noise, especially in robotics domain. To perform well even with such noise, we expect them to be able to detect outliers and discard them when…
In Online Learning to Rank (OLTR) the aim is to find an optimal ranking model by interacting with users. When learning from user behavior, systems must interact with users while simultaneously learning from those interactions. Unlike other…
We study the problem of serving randomly arriving and delay-sensitive traffic over a multi-channel communication system with time-varying channel states and unknown statistics. This problem deviates from the classical…
Maxmin-$\omega$ is a new threshold model, where each node in a network waits for the arrival of states from a fraction $\omega$ of neighborhood nodes before processing its own state, and subsequently transmitting it to downstream nodes.…
The current reinforcement learning framework focuses exclusively on performance, often at the expense of efficiency. In contrast, biological control achieves remarkable performance while also optimizing computational energy expenditure and…
The optimal (`equilibrium') macroscopic properties of an economy with $N$ industries endowed with different technologies, $P$ commodities and one consumer are derived in the limit $N\to\infty$ with $n=N/P$ fixed using the replica method.…
While research of reinforcement learning applied to financial markets predominantly concentrates on finding optimal behaviours, it is worth to realize that the reinforcement learning returns $G_t$ and state value functions themselves are of…
Who gains and who loses from a manipulable school-choice mechanism? Studying the outcomes of sincere and sophisticated students under the manipulable Boston Mechanism as compared with the strategy-proof Deferred Acceptance, we provide…
This study investigates the development of an optimal execution strategy through reinforcement learning, aiming to determine the most effective approach for traders to buy and sell inventory within a finite time horizon. Our proposed model…
We consider the sequential decision-making problem where the mean outcome is a non-linear function of the chosen action. Compared with the linear model, two curious phenomena arise in non-linear models: first, in addition to the "learning…
This paper studies social optimal control of mean field LQG (linear-quadratic-Gaussian) models with uncertainty. Specially, the uncertainty is represented by a uncertain drift which is common for all agents. A robust optimization approach…
We consider one buyer and one seller. For a bundle $(t,q)\in [0,\infty[\times [0,1]=\mathbb{Z}$, $q$ either refers to the wining probability of an object or a share of a good, and $t$ denotes the payment that the buyer makes. We define…
We develop a hyperparameter optimisation algorithm, Automated Budget Constrained Training (AutoBCT), which balances the quality of a model with the computational cost required to tune it. The relationship between hyperparameters, model…
In this paper we present a framework for risk-sensitive model predictive control (MPC) of linear systems affected by stochastic multiplicative uncertainty. Our key innovation is to consider a time-consistent, dynamic risk evaluation of the…
In recent years, the dominance of machine learning in stock market forecasting has been evident. While these models have shown decreasing prediction errors, their robustness across different datasets has been a concern. A successful stock…
A finite horizon optimal tracking problem is considered for linear dynamical systems subject to parametric uncertainties in the state-space matrices and exogenous disturbances. A suboptimal solution is proposed using a model predictive…
We study conditions for the existence of stable and group-strategy-proof mechanisms in a many-to-one matching model with contracts if students' preferences are monotone in contract terms. We show that "equivalence", properly defined, to a…
We investigate activities that have different periods of duration. We define the profit intensity as a measure of this economic category. The profit intensity in a repeated trading has a unique property of attaining its maximum at a fixed…