English
Related papers

Related papers: Aggregating time preferences with decreasing impat…

200 papers

Preference optimization is widely used to align large language models (LLMs) with human preferences. However, many margin-based methods also suppress the chosen response when they try to suppress the rejected one, and there is no general…

Machine Learning · Computer Science 2026-05-04 Wei Chen , Yubing Wu , Junmei Yang , Delu Zeng , Qibin Zhao , John Paisley , Min Chen , Zhou Wang

While Value Iteration (VI) is one of the most fundamental algorithms in Reinforcement Learning, its theoretical convergence guarantees still exhibit a persistent mismatch with empirical behavior. In the discounted-reward case, classical…

Machine Learning · Computer Science 2026-03-12 Arsenii Mustafin , Xinyi Sheng , Dominik Baumann

We study an optimal stopping problem under non-exponential discounting, where the state process is a multi-dimensional continuous strong Markov process. The discount function is taken to be log sub-additive, capturing decreasing impatience…

Mathematical Finance · Quantitative Finance 2021-07-14 Yu-Jui Huang , Zhenhua Wang

We consider an individual or household endowed with an initial capital and an income, modeled as a deterministic process with a continuous drift rate. At first, we model the discounting rate as the price of a zero-coupon bond at zero under…

Optimization and Control · Mathematics 2016-04-01 Julia Eisenberg

Preference aggregation is a fundamental problem in voting theory, in which public input rankings of a set of alternatives (called preferences) must be aggregated into a single preference that satisfies certain soundness properties. The…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-10-30 Kenan Wood , Hammurabi Mendes , Jonad Pulaj

The Difference in Difference (DiD) estimator is a popular estimator built on the "parallel trends" assumption, which is an assertion that the treatment group, absent treatment, would change "similarly" to the control group over time. To…

Methodology · Statistics 2024-02-09 Dae Woong Ham , Luke Miratrix

We study dynamic pricing where a seller repeatedly interacts with a strategic, non-myopic buyer who has a fixed private valuation and discounts future utility. Prior work focused exclusively on posted-price mechanisms, which only extract…

Computer Science and Game Theory · Computer Science 2026-04-28 Shiliang Zuo

We introduce a multiple criteria Bayesian preference learning framework incorporating behavioral cues for decision aiding. The framework integrates pairwise comparisons, response time, and attention duration to deepen insights into…

Applications · Statistics 2025-04-22 Jiaxuan Jiang , Jiapeng Liu , Miłosz Kadziński , Xiuwu Liao , Jingyu Dong

We investigate a value-maximizing problem incorporating a human behavior pattern: present-biased-ness, for a firm which navigates strategic decisions encompassing earning retention/payout and capital injection policies, within the framework…

Optimization and Control · Mathematics 2024-01-30 Kaixin Yan , Wenyuan Wang , Jinxia Zhu

The appropriate discount rate for evaluating policies is a critical consideration in economic decision-making. This paper presents a new model for calculating the derived discount rate for a society that includes different groups with…

Theoretical Economics · Economics 2025-02-11 Mahdi Mousavi , Mahdi Kohan Sefidi

Emphatic algorithms are temporal-difference learning algorithms that change their effective state distribution by selectively emphasizing and de-emphasizing their updates on different time steps. Recent works by Sutton, Mahmood and White…

Machine Learning · Computer Science 2015-07-07 A. Rupam Mahmood , Huizhen Yu , Martha White , Richard S. Sutton

Pricing decisions of companies require an understanding of the causal effect of a price change on the demand. When real-life pricing experiments are infeasible, data-driven decision-making must be based on alternative data sources such as…

Applications · Statistics 2024-07-03 Lauri Valkonen , Santtu Tikka , Jouni Helske , Juha Karvanen

We study a \emph{financial} version of the classic online problem of scheduling weighted packets with deadlines. The main novelty is that, while previous works assume packets have \emph{fixed} weights throughout their lifetime, this work…

Computer Science and Game Theory · Computer Science 2025-02-21 Yotam Gafni , Aviv Yaish

Cyclic coordinate descent is a classic optimization method that has witnessed a resurgence of interest in machine learning. Reasons for this include its simplicity, speed and stability, as well as its competitive performance on $\ell_1$…

Machine Learning · Computer Science 2015-03-17 Ankan Saha , Ambuj Tewari

This paper proposes an empirical model of dynamic discrete choice to allow for non-separable time preferences, generalizing the well-known Rust (1987) model. Under weak conditions, we show the existence of value functions and hence…

Econometrics · Economics 2024-06-13 Jay Lu , Yao Luo , Kota Saito , Yi Xin

A recent line of work, starting with Beigman and Vohra (2006) and Zadimoghaddam and Roth (2012), has addressed the problem of {\em learning} a utility function from revealed preference data. The goal here is to make use of past data…

Computer Science and Game Theory · Computer Science 2014-07-31 Maria-Florina Balcan , Amit Daniely , Ruta Mehta , Ruth Urner , Vijay V. Vazirani

User preference learning is generally a hard problem. Individual preferences are typically unknown even to users themselves, while the space of choices is infinite. Here we study user preference learning from information-theoretic…

Machine Learning · Computer Science 2023-11-27 Tanya Ignatenko , Kirill Kondrashov , Marco Cox , Bert de Vries

Motivated by applications to distributed optimization over networks and large-scale data processing in machine learning, we analyze the deterministic incremental aggregated gradient method for minimizing a finite sum of smooth functions…

Optimization and Control · Mathematics 2018-01-16 Mert Gurbuzbalaban , Asuman Ozdaglar , Pablo Parrilo

A firm that sells a non perishable product considers intertemporal price discrimination in the objective of maximizing its long-run average revenue. We consider a general model of patient customers with changing valuations. Arriving…

Optimization and Control · Mathematics 2020-02-17 Araman Victor , Fayad Bassam

This letter exposes a tight connection between the thermodynamic efficiency of information processing and predictive inference. A generalized lower bound on dissipation is derived for partially observable information engines which are…

Statistical Mechanics · Physics 2020-02-12 Susanne Still
‹ Prev 1 4 5 6 7 8 10 Next ›