中文
相关论文

相关论文: Stateful Posted Pricing with Vanishing Regret via …

200 篇论文

We study dynamic regret in federated online decision-making with stateful incurred costs under block-based synchronization and partial client participation. In this setting, sparse communication affects not only the pointwise update quality…

系统与控制 · 电气工程与系统科学 2026-05-18 Yiwei Liu , Luwei Yang , Shunbo Lei

As data marketplaces become increasingly central to the digital economy, it is crucial to design efficient pricing mechanisms that optimize revenue while ensuring fair and adaptive pricing. We introduce the Maximum Auction-to-Posted Price…

机器学习 · 统计学 2026-04-06 Yingqi Gao , Wenlu Xu , Jin J. Zhou , Hua Zhou , Yong Chen , Xiaowu Dai

Real-time bidding (RTB) is an important mechanism in online display advertising, where a proper bid for each page view plays an essential role for good marketing results. Budget constrained bidding is a typical scenario in RTB where the…

人工智能 · 计算机科学 2018-10-24 Di Wu , Xiujun Chen , Xun Yang , Hao Wang , Qing Tan , Xiaoxun Zhang , Jian Xu , Kun Gai

The agency problem emerges in today's large scale machine learning tasks, where the learners are unable to direct content creation or enforce data collection. In this work, we propose a theoretical framework for aligning economic interests…

机器学习 · 计算机科学 2024-07-03 Jibang Wu , Siyu Chen , Mengdi Wang , Huazheng Wang , Haifeng Xu

We study online decision making problems under resource constraints, where both reward and cost functions are drawn from distributions that may change adversarially over time. We focus on two canonical settings: $(i)$ online resource…

We investigate online Markov Decision Processes (MDPs) with adversarially changing loss functions and known transitions. We choose dynamic regret as the performance measure, defined as the performance difference between the learner and any…

机器学习 · 计算机科学 2022-08-29 Peng Zhao , Long-Fei Li , Zhi-Hua Zhou

We consider multiple parallel Markov decision processes (MDPs) coupled by global constraints, where the time varying objective and constraint functions can only be observed after the decision is made. Special attention is given to how well…

最优化与控制 · 数学 2017-09-12 Xiaohan Wei , Hao Yu , Michael J. Neely

We propose a novel change point detection approach for online learning control with full information feedback (state, disturbance, and cost feedback) for unknown time-varying dynamical systems. We show that our algorithm can achieve a…

系统与控制 · 电气工程与系统科学 2023-03-28 Deepan Muthirayan , Ruijie Du , Yanning Shen , Pramod P. Khargonekar

This paper studies the online optimal control problem with time-varying convex stage costs for a time-invariant linear dynamical system, where a finite lookahead window of accurate predictions of the stage costs are available at each time.…

最优化与控制 · 数学 2019-10-23 Yingying Li , Xin Chen , Na Li

Price responsiveness is a major feature of end use customers (EUCs) that participate in demand response (DR) programs, and has been conventionally modeled with static demand functions, which take the electricity price as the input and the…

机器学习 · 计算机科学 2020-06-09 Hanchen Xu , Hongbo Sun , Daniel Nikovski , Kitamura Shoichi , Kazuyuki Mori

We consider the problem of online dynamic mechanism design for sequential auctions in unknown environments, where the underlying market and, thus, the bidders' values vary over time as interactions between the seller and the bidders…

计算机科学与博弈论 · 计算机科学 2025-10-27 Vincent Leon , S. Rasoul Etesami

In this work we consider the online control of a known linear dynamic system with adversarial disturbance and adversarial controller cost. The goal in online control is to minimize the regret, defined as the difference between cumulative…

最优化与控制 · 数学 2021-10-15 Deepan Muthirayan , Jianjun Yuan , Pramod P. Khargonekar

This paper proposes a novel approach to construct data-driven online solutions to optimization problems (P) subject to a class of distributionally uncertain dynamical systems. The introduced framework allows for the simultaneous learning of…

系统与控制 · 电气工程与系统科学 2024-07-23 Dan Li , Dariush Fooladivanda , Sonia Martinez

In this paper we present an end-to-end framework for addressing the problem of dynamic pricing (DP) on E-commerce platform using methods based on deep reinforcement learning (DRL). By using four groups of different business data to…

机器学习 · 计算机科学 2021-09-01 Jiaxi Liu , Yidong Zhang , Xiaoqing Wang , Yuming Deng , Xingyu Wu

In learning theory, the performance of an online policy is commonly measured in terms of the static regret metric, which compares the cumulative loss of an online policy to that of an optimal benchmark in hindsight. In the definition of…

信息论 · 计算机科学 2022-08-23 Ativ Joshi , Abhishek Sinha

We consider an agent who is involved in a Markov decision process and receives a vector of outcomes every round. Her objective is to maximize a global concave reward function on the average vectorial outcome. The problem models applications…

机器学习 · 计算机科学 2019-05-17 Wang Chi Cheung

Recently, several universal methods have been proposed for online convex optimization which can handle convex, strongly convex and exponentially concave cost functions simultaneously. However, most of these algorithms have been designed…

机器学习 · 计算机科学 2023-02-14 Arnold Salas

Dynamic resource allocation problems are ubiquitous, arising in inventory management, order fulfillment, online advertising, and other applications. We initially focus on one of the simplest models of online resource allocation: the…

概率论 · 数学 2025-06-04 Omar Besbes , Yash Kanoria , Akshit Kumar

The online Markov decision process (MDP) is a generalization of the classical Markov decision process that incorporates changing reward functions. In this paper, we propose practical online MDP algorithms with policy iteration and…

机器学习 · 计算机科学 2015-10-16 Yao Ma , Hao Zhang , Masashi Sugiyama

Traditional reinforcement learning usually assumes either episodic interactions with resets or continuous operation to minimize average or cumulative loss. While episodic settings have many theoretical results, resets are often unrealistic…

最优化与控制 · 数学 2026-01-13 Bianca Marin Moreno , Margaux Brégère , Pierre Gaillard , Nadia Oudjane