中文
相关论文

相关论文: Close the Gaps: A Learning-while-Doing Algorithm f…

200 篇论文

We consider the problem of \emph{optimal matching with queues} in dynamic systems and investigate the value-of-information. In such systems, the operators match tasks and resources stored in queues, with the objective of maximizing the…

最优化与控制 · 数学 2015-03-30 Longbo Huang

The optimal objective is a fundamental aspect of reinforcement learning (RL), as it determines how policies are evaluated and optimized. While total return maximization is the ideal objective in RL, discounted return maximization is the…

机器学习 · 计算机科学 2025-03-19 Shuyu Yin , Fei Wen , Peilin Liu , Tao Luo

In online marketplaces, customers have access to hundreds of reviews for a single product. Buyers often use reviews from other customers that share their type -- such as height for clothing, skin type for skincare products, and location for…

计算机科学与博弈论 · 计算机科学 2023-09-12 Wenshuo Guo , Nika Haghtalab , Kirthevasan Kandasamy , Ellen Vitercik

Internet advertisers (buyers) repeatedly procure ad impressions from ad platforms (sellers) with the aim to maximize total conversion (i.e. ad value) while respecting both budget and return-on-investment (ROI) constraints for efficient…

计算机科学与博弈论 · 计算机科学 2023-02-08 Negin Golrezaei , Patrick Jaillet , Jason Cheuk Nam Liang , Vahab Mirrokni

As data marketplaces become increasingly central to the digital economy, it is crucial to design efficient pricing mechanisms that optimize revenue while ensuring fair and adaptive pricing. We introduce the Maximum Auction-to-Posted Price…

机器学习 · 统计学 2026-04-06 Yingqi Gao , Wenlu Xu , Jin J. Zhou , Hua Zhou , Yong Chen , Xiaowu Dai

I consider the optimal hourly (or per-unit-time in general) pricing problem faced by a freelance worker (or a service provider) on an on-demand service platform. Service requests arriving while the worker is busy are lost forever. Thus, the…

计算机科学与博弈论 · 计算机科学 2019-07-09 Vijay Kamble

We study an online market-making problem in which a learner sequentially posts bid and ask prices for a single asset while interacting with traders holding private valuations. Unlike existing online learning formulations that assume fully…

机器学习 · 计算机科学 2026-05-20 Davide Maran , Marcello Restelli

We consider a two-product inventory system with independent Poisson demands, limited joint storage capacity and partial demand substitution. Replenishment is performed simultaneously for both products and the replenishment time may be fixed…

最优化与控制 · 数学 2015-10-20 Apostolos N. Burnetas , Odysseas Kanavetas

Optimal execution is a sequential decision-making problem for cost-saving in algorithmic trading. Studies have found that reinforcement learning (RL) can help decide the order-splitting sizes. However, a problem remains unsolved: how to…

交易与市场微观结构 · 定量金融 2022-07-25 Feiyang Pan , Tongzhe Zhang , Ling Luo , Jia He , Shuoling Liu

Most recent research in network revenue management incorporates choice behavior that models the customers' buying logic. These models are consequently more complex to solve, but they return a more robust policy that usually generates better…

最优化与控制 · 数学 2019-11-05 Thibault Barbier , Miguel Anjos , Fabien Cirinei , Gilles Savard

In the Learning to Price setting, a seller posts prices over time with the goal of maximizing revenue while learning the buyer's valuation. This problem is very well understood when values are stationary (fixed or iid). Here we study the…

计算机科学与博弈论 · 计算机科学 2021-06-10 Renato Paes Leme , Balasubramanian Sivan , Yifeng Teng , Pratik Worah

Learning about many things can provide numerous benefits to a reinforcement learning system. For example, learning many auxiliary value functions, in addition to optimizing the environmental reward, appears to improve both exploration and…

机器学习 · 计算机科学 2020-08-25 Cam Linke , Nadia M. Ady , Martha White , Thomas Degris , Adam White

We study a continuous-time, infinite-horizon dynamic bipartite matching problem. Suppliers arrive according to a Poisson process; while waiting, they may abandon the queue at a uniform rate. Customers on the other hand must be matched upon…

数据结构与算法 · 计算机科学 2025-06-03 Alireza AmaniHamedani , Ali Aouad , Amin Saberi

The Joint Replenishment Problem (JRP) is a fundamental optimization problem in supply-chain management, concerned with optimizing the flow of goods from a supplier to retailers. Over time, in response to demands at the retailers, the…

We study the problem of designing posted-price mechanisms in order to sell a single unit of a single item within a finite period of time. Motivated by real-world problems, such as, e.g., long-term rental of rooms and apartments, we assume…

计算机科学与博弈论 · 计算机科学 2020-12-11 Giulia Romano , Gianluca Tartaglia , Alberto Marchesi , Nicola Gatti

We study non-stationary single-item, periodic-review inventory control problems in which the demand distribution is unknown and may change over time. We analyze how demand non-stationarity affects learning performance across inventory…

最优化与控制 · 数学 2026-02-06 Nele H. Amiri , Sean R. Sinclair , Maximiliano Udenio

Peer prediction mechanisms are often adopted to elicit truthful contributions from crowd workers when no ground-truth verification is available. Recently, mechanisms of this type have been developed to incentivize effort exertion, in…

计算机科学与博弈论 · 计算机科学 2016-12-05 Yang Liu , Yiling Chen

Online advertising platforms use automated auctions to connect advertisers with potential customers, requiring effective bidding strategies to maximize profits. Accurate ad impact estimation requires considering three key factors: delayed…

机器学习 · 计算机科学 2025-10-24 Yuwei Cheng , Zifeng Zhao , Haifeng Xu

In business and marketing, analyzing the reasons behind buying is a fundamental step towards understanding consumer behaviors, shaping business strategies, and predicting market outcomes. Prior research on purchase reason has relied on…

信息检索 · 计算机科学 2024-11-19 Tao Chen , Siqi Zuo , Cheng Li , Mingyang Zhang , Qiaozhu Mei , Michael Bendersky

In modern e-commerce and service operations, firms must jointly manage inventory replenishment and real-time order fulfillment to maximize profit under demand uncertainty. While each component has been studied extensively in isolation,…

最优化与控制 · 数学 2026-03-05 Zi Ling , Jiashuo Jiang , Linwei Xin