中文
相关论文

相关论文: Bid Optimization using Maximum Entropy Reinforceme…

200 篇论文

Bid shading plays a crucial role in Real-Time Bidding (RTB) by adaptively adjusting the bid to avoid advertisers overspending. Existing mainstream two-stage methods, which first model bid landscapes and then optimize surplus using…

计算机科学与博弈论 · 计算机科学 2026-04-30 Yinqiu Huang , Hao Ma , Wenshuai Chen , Zongwei Wang , Shuli Wang , Yongqiang Zhang , Xue Wei , Yinhua Zhu , Haitao Wang , Xingxing Wang

While most approaches to the problem of Inverse Reinforcement Learning (IRL) focus on estimating a reward function that best explains an expert agent's policy or demonstrated behavior on a control task, it is often the case that such…

机器学习 · 计算机科学 2020-05-01 Dexter R. R. Scobee , S. Shankar Sastry

Auto-bidding plays an important role in online advertising and has become a crucial tool for advertisers and advertising platforms to meet their performance objectives and optimize the efficiency of ad delivery. Advertisers employing…

计算机科学与博弈论 · 计算机科学 2020-12-07 Bin Li , Xiao Yang , Daren Sun , Zhi Ji , Zhen Jiang , Cong Han , Dong Hao

Display advertising is an important online advertising type where banner advertisements (shortly ad) on websites are usually measured by how many times they are viewed by online users. There are two major channels to sell ad views. They can…

计算机科学与博弈论 · 计算机科学 2017-01-20 Bowei Chen

Recently the online advertising market has exhibited a gradual shift from second-price auctions to first-price auctions. Although there has been a line of works concerning online bidding strategies in first-price auctions, it still remains…

计算机科学与博弈论 · 计算机科学 2022-05-31 Rui Ai , Chang Wang , Chenchen Li , Jinshan Zhang , Wenhan Huang , Xiaotie Deng

To overcome the curses of dimensionality and modeling of Dynamic Programming (DP) methods to solve Markov Decision Process (MDP) problems, Reinforcement Learning (RL) methods are adopted in practice. Contrary to traditional RL algorithms…

机器学习 · 计算机科学 2021-08-24 Arghyadip Roy , Vivek Borkar , Abhay Karandikar , Prasanna Chaporkar

Online retailers often use third-party demand-side-platforms (DSPs) to conduct offsite advertising and reach shoppers across the Internet on behalf of their advertisers. The process involves the retailer participating in instant auctions…

系统与控制 · 电气工程与系统科学 2023-10-26 Hangjian Li , Dong Xu , Konstantin Shmakov , Kuang-Chih Lee , Wei Shen

The Empirical Revenue Maximization (ERM) is one of the most important price learning algorithms in auction design: as the literature shows it can learn approximately optimal reserve prices for revenue-maximizing auctioneers in both repeated…

计算机科学与博弈论 · 计算机科学 2020-10-13 Xiaotie Deng , Ron Lavi , Tao Lin , Qi Qi , Wenwei Wang , Xiang Yan

Learning to bid in repeated first-price auctions is a fundamental problem at the interface of game theory and machine learning, which has seen a recent surge in interest due to the transition of display advertising to first-price auctions.…

计算机科学与博弈论 · 计算机科学 2024-07-09 Rachitesh Kumar , Jon Schneider , Balasubramanian Sivan

Abstract In this work, we build two environments, namely the modified QLBS and RLOP models, from a mathematics perspective which enables RL methods in option pricing through replicating by portfolio. We implement the environment…

证券定价 · 定量金融 2022-05-12 Ziheng Chen

We study the problem of allocating impressions to sellers in e-commerce websites, such as Amazon, eBay or Taobao, aiming to maximize the total revenue generated by the platform. We employ a general framework of reinforcement mechanism…

多智能体系统 · 计算机科学 2018-02-28 Qingpeng Cai , Aris Filos-Ratsikas , Pingzhong Tang , Yiwei Zhang

This research focuses on the bid optimization problem in the real-time bidding setting for online display advertisements, where an advertiser, or the advertiser's agent, has access to the features of the website visitor and the type of ad…

机器学习 · 计算机科学 2022-10-31 Rui Fan , Erick Delage

Safe Reinforcement Learning (RL) plays an important role in applying RL algorithms to safety-critical real-world applications, addressing the trade-off between maximizing rewards and adhering to safety constraints. This work introduces a…

机器人学 · 计算机科学 2024-07-16 Fan Yang , Wenxuan Zhou , Zuxin Liu , Ding Zhao , David Held

To increase brand awareness, many advertisers conclude contracts with advertising platforms to purchase traffic and then deliver advertisements to target audiences. In a whole delivery period, advertisers usually desire a certain impression…

信息检索 · 计算机科学 2024-06-18 Penghui Wei , Yongqiang Chen , Shaoguo Liu , Liang Wang , Bo Zheng

In the realm of online advertising, advertisers partake in ad auctions to obtain advertising slots, frequently taking advantage of auto-bidding tools provided by demand-side platforms. To improve the automation of these bidding systems, we…

机器学习 · 计算机科学 2025-06-30 Hao Jiang , Yongxiang Tang , Yanxiang Zeng , Pengjia Yuan , Yanhua Cheng , Teng Sha , Xialong Liu , Peng Jiang

Maximum entropy reinforcement learning integrates exploration into policy learning by providing additional intrinsic rewards proportional to the entropy of some distribution. In this paper, we propose a novel approach in which the intrinsic…

机器学习 · 计算机科学 2025-09-30 Adrien Bolland , Gaspard Lambrechts , Damien Ernst

Reinforcement learning (RL) agents have traditionally been tasked with maximizing the value function of a Markov decision process (MDP), either in continuous settings, with fixed discount factor $\gamma < 1$, or in episodic settings, with…

机器学习 · 计算机科学 2019-02-11 Silviu Pitis

This paper develops a novel multi-agent reinforcement learning (MARL) framework for reinsurance treaty bidding, addressing long-standing inefficiencies in traditional broker-mediated placement processes. We pose the core research question:…

人工智能 · 计算机科学 2026-03-24 Stella C. Dong , James R. Finlay

The emergence of price comparison websites (PCWs) has presented insurers with unique challenges in formulating effective pricing strategies. Operating on PCWs requires insurers to strike a delicate balance between competitive premiums and…

In this thesis, we research learning algorithms for optimal decision making in two different contexts, Reinforcement Learning in Part I and Auction Design in Part II. Reinforcement learning (RL) is an area of machine learning that is…

机器学习 · 计算机科学 2022-10-07 Jad Rahme