English
Related papers

Related papers: Reading policies for joins: An asymptotic analysis

200 papers

We prove that the consumption functions in optimal savings problems are asymptotically linear if the marginal utility is regularly varying. We also analytically characterize the asymptotic marginal propensities to consume (MPCs) out of…

General Economics · Economics 2021-10-26 Qingyin Ma , Alexis Akira Toda

This paper reexamines Abadie and Imbens (2016)'s work on propensity score matching for average treatment effect estimation. We explore the asymptotic behavior of these estimators when the number of nearest neighbors, $M$, grows with the…

Statistics Theory · Mathematics 2023-11-16 Yihui He , Fang Han

Policy learning can be used to extract individualized treatment regimes from observational data in healthcare, civics, e-commerce, and beyond. One big hurdle to policy learning is a commonplace lack of overlap in the data for different…

Machine Learning · Statistics 2020-12-04 Nathan Kallus

For effective decision support in scenarios with conflicting objectives, sets of potentially optimal solutions can be presented to the decision maker. We explore both what policies these sets should contain and how such sets can be computed…

Artificial Intelligence · Computer Science 2023-07-19 Willem Röpke , Conor F. Hayes , Patrick Mannion , Enda Howley , Ann Nowé , Diederik M. Roijers

Service platforms must determine rules for matching heterogeneous demand (customers) and supply (workers) that arrive randomly over time and may be lost if forced to wait too long for a match. Our objective is to maximize the cumulative…

Optimization and Control · Mathematics 2023-12-19 Angelos Aveklouris , Levi DeValve , Maximiliano Stock , Amy R. Ward

We study and compare three estimators of a discrete monotone distribution: (a) the (raw) empirical estimator; (b) the "method of rearrangements" estimator; and (c) the maximum likelihood estimator. We show that the maximum likelihood…

Statistics Theory · Mathematics 2009-10-20 Hanna K. Jankowski , Jon A. Wellner

We consider a unified framework of sequential change-point detection and hypothesis testing modeled by means of hidden Markov chains. One observes a sequence of random variables whose distributions are functionals of a hidden Markov chain.…

Optimization and Control · Mathematics 2013-12-13 Savas Dayanik , Kazutoshi Yamazaki

Maximum entropy reinforcement learning integrates exploration into policy learning by providing additional intrinsic rewards proportional to the entropy of some distribution. In this paper, we propose a novel approach in which the intrinsic…

Machine Learning · Computer Science 2025-09-30 Adrien Bolland , Gaspard Lambrechts , Damien Ernst

Let $\alpha_n(\cdot)=P\bigl(X_{n+1}\in\cdot\mid X_1,\ldots,X_n\bigr)$ be the predictive distributions of a sequence $(X_1,X_2,\ldots)$ of $p$-dimensional random vectors. Suppose $$\alpha_n= \mathcal{N} _p (M_n,Q_n)$$ where…

Statistics Theory · Mathematics 2024-09-17 Samuele Garelli , Fabrizio Leisen , Luca Pratelli , Pietro Rigo

We find a two term asymptotic expansion for the optimal expected value of a sequentially selected monotone subsequence from a random permutation of length n. A striking feature of this expansion is that tells us that the expected value of…

Probability · Mathematics 2015-09-16 Peichao Peng , J. Michael Steele

We consider n agents located on the vertices of a connected graph. Each agent v receives a signal X_v(0)~N(s, 1) where s is an unknown quantity. A natural iterative way of estimating s is to perform the following procedure. At iteration t +…

Statistics Theory · Mathematics 2010-07-13 Elchanan Mossel , Omer Tamuz

We introduce a simple time-triggered protocol to achieve communication-efficient non-Bayesian learning over a network. Specifically, we consider a scenario where a group of agents interact over a graph with the aim of discerning the true…

Systems and Control · Electrical Eng. & Systems 2019-09-05 Aritra Mitra , John A. Richards , Shreyas Sundaram

We consider a learning system based on the conventional multiplicative weight (MW) rule that combines experts' advice to predict a sequence of true outcomes. It is assumed that one of the experts is malicious and aims to impose the maximum…

Machine Learning · Computer Science 2020-09-21 S. Rasoul Etesami , Negar Kiyavash , Vincent Leon , H. Vincent Poor

The analysis of strings of $n$ random variables with geometric distribution has recently attracted renewed interest: Archibald et al. consider the number of distinct adjacent pairs in geometrically distributed words. They obtain the…

Probability · Mathematics 2024-02-14 Guy Louchard , Werner Schachinger , Mark Daniel Ward

Exploration and adaptation to new tasks in a transfer learning setup is a central challenge in reinforcement learning. In this work, we build on the idea of modeling a distribution over policies in a Bayesian deep reinforcement learning…

Machine Learning · Computer Science 2019-06-11 Disha Shrivastava , Eeshan Gunesh Dhekane , Riashat Islam

The random greedy algorithm for finding a maximal independent set in a graph constructs a maximal independent set by inspecting the graph's vertices in a random order, adding the current vertex to the independent set if it is not adjacent…

Combinatorics · Mathematics 2023-09-28 Michael Krivelevich , Tamás Mészáros , Peleg Michaeli , Clara Shikhelman

The asymptotic behavior of the stochastic gradient algorithm with a biased gradient estimator is analyzed. Relying on arguments based on the dynamic system theory (chain-recurrence) and the differential geometry (Yomdin theorem and…

Statistics Theory · Mathematics 2017-09-04 Vladislav B. Tadic , Arnaud Doucet

This paper studies optimal linear power control for battery-limited energy harvesting communications. It provides a systematic analysis of linear power control policies, covering the greedy and fixed-fraction policies as special cases.…

Information Theory · Computer Science 2026-02-25 Hafez M. Garmaroudi , Zikai Dou , Shengtian Yang , Jun Chen

This paper introduces the first asymptotically optimal strategy for a multi armed bandit (MAB) model under side constraints. The side constraints model situations in which bandit activations are limited by the availability of certain…

Machine Learning · Statistics 2025-02-10 Apostolos N. Burnetas , Odysseas Kanavetas , Michael N. Katehakis

We identify and study the phenomenon of policy churn, that is, the rapid change of the greedy policy in value-based reinforcement learning. Policy churn operates at a surprisingly rapid pace, changing the greedy action in a large fraction…

Machine Learning · Computer Science 2022-10-24 Tom Schaul , André Barreto , John Quan , Georg Ostrovski