中文
相关论文

相关论文: Breaking Determinism: Stochastic Modeling for Reli…

200 篇论文

We present a deterministic exploration mechanism for sponsored search auctions, which enables the auctioneer to learn the relevance scores of advertisers, and allows advertisers to estimate the true value of clicks generated at the auction…

计算机科学与博弈论 · 计算机科学 2011-11-10 Sudhir Kumar Singh , Vwani P. Roychowdhury , Milan Bradonjić , Behnam A. Rezaei

In programmatic advertising, ad slots are usually sold using second-price (SP) auctions in real-time. The highest bidding advertiser wins but pays only the second-highest bid (known as the winning price). In SP, for a single item, the…

机器学习 · 计算机科学 2020-01-22 Aritra Ghosh , Saayan Mitra , Somdeb Sarkhel , Jason Xie , Gang Wu , Viswanathan Swaminathan

Efficient large-scale network allocation requires data-driven pricing mechanisms that internalize the stochastic and non-linear dynamics of user behavior. We move beyond the classic fully strategic agents to study oblivious users (agents…

数值分析 · 数学 2026-05-28 Yixuan Li , Andersen Ang , Sebastian Stein

Optimizing survival outcomes, such as patient survival or customer retention, is a critical objective in data-driven decision-making. Off-Policy Evaluation~(OPE) provides a powerful framework for assessing such decision-making policies…

统计方法学 · 统计学 2026-03-25 Kohsuke Kubota , Mitsuhiro Takahashi , Yuta Saito

We consider off-policy evaluation (OPE), which evaluates the performance of a new policy from observed data collected from previous experiments, without requiring the execution of the new policy. This finds important applications in areas…

机器学习 · 计算机科学 2020-08-18 Yihao Feng , Tongzheng Ren , Ziyang Tang , Qiang Liu

In cost-per-click (CPC) or cost-per-impression (CPM) advertising campaigns, advertisers always run the risk of spending the budget without getting enough conversions. Moreover, the bidding on advertising inventory has few connections with…

信息检索 · 计算机科学 2022-12-29 Deguang Kong , Konstantin Shmakov , Jian Yang

Online retailers often use third-party demand-side-platforms (DSPs) to conduct offsite advertising and reach shoppers across the Internet on behalf of their advertisers. The process involves the retailer participating in instant auctions…

系统与控制 · 电气工程与系统科学 2023-10-26 Hangjian Li , Dong Xu , Konstantin Shmakov , Kuang-Chih Lee , Wei Shen

Recent advances in machine learning have spurred significant interest in learning-augmented algorithms, particularly for online optimization. A growing body of work has studied online bidding in this framework, aiming to characterize the…

数据结构与算法 · 计算机科学 2026-05-11 Changyeol Lee , Dahoon Lee , Jongseo Lee , Yongho Shin , Changki Yun

Offline reinforcement learning enables agents to leverage large pre-collected datasets of environment transitions to learn control policies, circumventing the need for potentially expensive or unsafe online data collection. Significant…

机器学习 · 计算机科学 2022-03-17 Cong Lu , Philip J. Ball , Jack Parker-Holder , Michael A. Osborne , Stephen J. Roberts

We study distributional off-policy evaluation (OPE), of which the goal is to learn the distribution of the return for a target policy using offline data generated by a different policy. The theoretical foundation of many existing work…

机器学习 · 统计学 2025-03-13 Sungee Hong , Zhengling Qi , Raymond K. W. Wong

The standard framework of online bidding algorithm design assumes that the seller commits himself to faithfully implementing the rules of the adopted auction. However, the seller may attempt to cheat in execution to increase his revenue if…

计算机科学与博弈论 · 计算机科学 2023-11-28 Qian Wang , Xuanzhi Xia , Zongjun Yang , Xiaotie Deng , Yuqing Kong , Zhilin Zhang , Liang Wang , Chuan Yu , Jian Xu , Bo Zheng

In recommender systems (RecSys) and real-time bidding (RTB) for online advertisements, we often try to optimize sequential decision making using bandit and reinforcement learning (RL) techniques. In these applications, offline reinforcement…

机器学习 · 计算机科学 2021-09-20 Haruka Kiyohara , Kosuke Kawakami , Yuta Saito

Predicting click and conversion probabilities when bidding on ad exchanges is at the core of the programmatic advertising industry. Two separated lines of previous works respectively address i) the prediction of user conversion probability…

机器学习 · 统计学 2017-07-24 Eustache Diemert , Julien Meynet , Pierre Galland , Damien Lefortier

In display advertising, a small group of sellers and bidders face each other in up to 10 12 auctions a day. In this context, revenue maximisation via monopoly price learning is a high-value problem for sellers. By nature, these auctions are…

机器学习 · 计算机科学 2020-10-21 Lorenzo Croissant , Marc Abeille , Clément Calauzènes

Automated bidding, an emerging intelligent decision making paradigm powered by machine learning, has become popular in online advertising. Advertisers in automated bidding evaluate the cumulative utilities and have private financial…

计算机科学与博弈论 · 计算机科学 2023-08-22 Yidan Xing , Zhilin Zhang , Zhenzhe Zheng , Chuan Yu , Jian Xu , Fan Wu , Guihai Chen

A/B testing is ubiquitous within the machine learning and data science operations of internet companies. Generically, the idea is to perform a statistical test of the hypothesis that a new feature is better than the existing platform---for…

统计理论 · 数学 2017-10-11 David Goldberg , James E. Johndrow

In this work, we consider the problem of estimating a behaviour policy for use in Off-Policy Policy Evaluation (OPE) when the true behaviour policy is unknown. Via a series of empirical studies, we demonstrate how accurate OPE is strongly…

Online advertising has become one of the most successful business models of the internet era. Impression opportunities are typically allocated through real-time auctions, where advertisers bid to secure advertisement slots. Deciding the…

机器学习 · 计算机科学 2025-05-20 Alberto Silvio Chiappa , Briti Gangopadhyay , Zhao Wang , Shingo Takamatsu

We study the problem of off-policy evaluation (OPE) for episodic Partially Observable Markov Decision Processes (POMDPs) with continuous states. Motivated by the recently proposed proximal causal inference framework, we develop a…

机器学习 · 统计学 2022-10-18 Rui Miao , Zhengling Qi , Xiaoke Zhang

We design online algorithms for the fair allocation of public goods to a set of $N$ agents over a sequence of $T$ rounds and focus on improving their performance using predictions. In the basic model, a public good arrives in each round,…

计算机科学与博弈论 · 计算机科学 2022-10-03 Siddhartha Banerjee , Vasilis Gkatzelis , Safwan Hossain , Billy Jin , Evi Micha , Nisarg Shah