English
Related papers

Related papers: Parameter-Adaptive Dynamic Pricing

200 papers

Contextual multi-armed bandit (MAB) achieves cutting-edge performance on a variety of problems. When it comes to real-world scenarios such as recommendation system and online advertising, however, it is essential to consider the resource…

Machine Learning · Computer Science 2020-04-07 Mengyue Yang , Qingyang Li , Zhiwei Qin , Jieping Ye

We develop algorithms for online linear regression which achieve optimal static and dynamic regret guarantees \emph{even in the complete absence of prior knowledge}. We present a novel analysis showing that a discounted variant of the…

Machine Learning · Computer Science 2024-05-30 Andrew Jacobsen , Ashok Cutkosky

Real-time bidding (RTB) has become a major paradigm of display advertising. Each ad impression generated from a user visit is auctioned in real time, where demand-side platform (DSP) automatically provides bid price usually relying on the…

Information Retrieval · Computer Science 2022-12-26 Zhimeng Jiang , Kaixiong Zhou , Mi Zhang , Rui Chen , Xia Hu , Soo-Hyun Choi

We formulate a multi-armed bandit (MAB) approach to choosing expert policies online in Markov decision processes (MDPs). Given a set of expert policies trained on a state and action space, the goal is to maximize the cumulative reward of…

Systems and Control · Computer Science 2017-07-19 Eric Mazumdar , Roy Dong , Vicenç Rúbies Royo , Claire Tomlin , S. Shankar Sastry

Prompt tuning has emerged as a key technique for adapting large pre-trained Decision Transformers (DTs) in offline Reinforcement Learning (RL), particularly in multi-task and few-shot settings. The Prompting Decision Transformer (PDT)…

Machine Learning · Computer Science 2025-10-02 Finn Rietz , Oleg Smirnov , Sara Karimi , Lele Cao

The rapid advancement in large language models (LLMs) has brought forth a diverse range of models with varying capabilities that excel in different tasks and domains. However, selecting the optimal LLM for user queries often involves a…

Machine Learning · Computer Science 2025-02-06 Yang Li

We study a nonparametric contextual bandit problem where the expected reward functions belong to a H\"older class with smoothness parameter $\beta$. We show how this interpolates between two extremes that were previously studied in…

Machine Learning · Statistics 2020-09-14 Yichun Hu , Nathan Kallus , Xiaojie Mao

We study bandit convex optimization methods that adapt to the norm of the comparator, a topic that has only been studied before for its full-information counterpart. Specifically, we develop convex bandit algorithms with regret bounds that…

Machine Learning · Computer Science 2020-07-17 Dirk van der Hoeven , Ashok Cutkosky , Haipeng Luo

This work explores a novel approach for adaptive, differentiable parametrization of large-scale non-stationary random fields. Coupled with any gradient-based algorithm, the method can be applied to variety of optimization problems,…

Optimization and Control · Mathematics 2019-03-19 Andrei Mukhin , Aleksey Khlyupin

This paper addresses the estimation of a time- varying parameter in a network. A group of agents sequentially receive noisy signals about the parameter (or moving target), which does not follow any particular dynamics. The parameter is not…

Optimization and Control · Mathematics 2016-03-03 Shahin Shahrampour , Alexander Rakhlin , Ali Jadbabaie

We consider the problem of model selection for the general stochastic contextual bandits under the realizability assumption. We propose a successive refinement based algorithm called Adaptive Contextual Bandit ({\ttfamily ACB}), that works…

Machine Learning · Statistics 2023-07-21 Avishek Ghosh , Abishek Sankararaman , Kannan Ramchandran

In this paper, by leveraging abundant observational transaction data, we propose a novel data-driven and interpretable pricing approach for markdowns, consisting of counterfactual prediction and multi-period price optimization. Firstly, we…

Artificial Intelligence · Computer Science 2021-05-20 Junhao Hua , Ling Yan , Huan Xu , Cheng Yang

We study offline dynamic pricing when historical data provide incomplete coverage of the price space such that some candidate prices, including the optimal one, may be entirely unobserved. This setting is common in practice and is…

Machine Learning · Statistics 2026-05-25 Zeyu Bian , Lan Wang , Zhengling Qi

We propose algorithms for online principal component analysis (PCA) and variance minimization for adaptive settings. Previous literature has focused on upper bounding the static adversarial regret, whose comparator is the optimal fixed…

Machine Learning · Computer Science 2019-05-14 Jianjun Yuan , Andrew Lamperski

This paper explores a new form of the linear bandit problem in which the algorithm receives the usual stochastic rewards as well as stochastic feedback about which features are relevant to the rewards, the latter feedback being the novel…

Machine Learning · Computer Science 2019-03-13 Urvashi Oswal , Aniruddha Bhargava , Robert Nowak

We consider a stochastic inventory control problem under censored demands, lost sales, and positive lead times. This is a fundamental problem in inventory management, with significant literature establishing near-optimality of a simple…

Machine Learning · Computer Science 2019-05-14 Shipra Agrawal , Randy Jia

We study a class of nested path problems, in which every path-based variable can be decomposed into a sequence of subpaths. Subpaths must satisfy local resources, while paths must satisfy additional global resources. This paper develops a…

Optimization and Control · Mathematics 2026-05-28 Bart van Rossum , Rolf van Lieshout , Alexandre Jacquillat

In this paper we present a theoretical framework for determining dynamic ask and bid prices of derivatives using the theory of dynamic coherent acceptability indices in discrete time. We prove a version of the First Fundamental Theorem of…

Risk Management · Quantitative Finance 2013-06-13 Tomasz R. Bielecki , Igor Cialenco , Ismail Iyigunler , Rodrigo Rodriguez

Personalized pricing, which involves tailoring prices based on individual characteristics, is commonly used by firms to implement a consumer-specific pricing policy. In this process, buyers can also strategically manipulate their feature…

Machine Learning · Statistics 2024-06-27 Pangpang Liu , Zhuoran Yang , Zhaoran Wang , Will Wei Sun

In a low-rank linear bandit problem, the reward of an action (represented by a matrix of size $d_1 \times d_2$) is the inner product between the action and an unknown low-rank matrix $\Theta^*$. We propose an algorithm based on a novel…

Machine Learning · Statistics 2020-10-20 Yangyi Lu , Amirhossein Meisami , Ambuj Tewari