中文
相关论文

相关论文: Bid Optimization using Maximum Entropy Reinforceme…

200 篇论文

Online advertising has become one of the most successful business models of the internet era. Impression opportunities are typically allocated through real-time auctions, where advertisers bid to secure advertisement slots. Deciding the…

机器学习 · 计算机科学 2025-05-20 Alberto Silvio Chiappa , Briti Gangopadhyay , Zhao Wang , Shingo Takamatsu

In recommender systems (RecSys) and real-time bidding (RTB) for online advertisements, we often try to optimize sequential decision making using bandit and reinforcement learning (RL) techniques. In these applications, offline reinforcement…

机器学习 · 计算机科学 2021-09-20 Haruka Kiyohara , Kosuke Kawakami , Yuta Saito

In a day-ahead market, energy buyers and sellers submit their bids for a particular future time, including the amount of energy they wish to buy or sell and the price they are prepared to pay or receive. However, the dynamic for forming the…

最优化与控制 · 数学 2024-11-26 Luca Di Persio , Matteo Garbelli , Luca M. Giordano

Real-Time Bidding is a new Internet advertising system that has become very popular in recent years. This system works like a global auction where advertisers bid to display their impressions in the publishers' ad slots. The most popular…

计算机科学与博弈论 · 计算机科学 2020-10-26 Luis Miralles-Pechuán , Fernando Jiménez , José Manuel García

Many real-world auctions are dynamic processes, in which bidders interact and report information over multiple rounds with the auctioneer. The sequential decision making aspect paired with imperfect information renders analyzing the…

计算机科学与博弈论 · 计算机科学 2023-12-21 Vinzenz Thoma , Michael Curry , Niao He , Sven Seuken

With the recent prevalence of Reinforcement Learning (RL), there have been tremendous interests in utilizing RL for online advertising in recommendation platforms (e.g., e-commerce and news feed sites). However, most RL-based advertising…

信息检索 · 计算机科学 2021-05-06 Xiangyu Zhao , Changsheng Gu , Haoshenglun Zhang , Xiwang Yang , Xiaobing Liu , Jiliang Tang , Hui Liu

With the proliferation of advanced metering infrastructure (AMI), more real-time data is available to electric utilities and consumers. Such high volumes of data facilitate innovative electricity rate structures beyond flat-rate and…

系统与控制 · 电气工程与系统科学 2021-11-23 Eli Brock , Lauren Bruckstein , Patrick Connor , Sabrina Nguyen , Robert Kerestes , Mai Abdelhakim

Online advertising platforms use automated auctions to connect advertisers with potential customers, requiring effective bidding strategies to maximize profits. Accurate ad impact estimation requires considering three key factors: delayed…

机器学习 · 计算机科学 2025-10-24 Yuwei Cheng , Zifeng Zhao , Haifeng Xu

Managing millions of digital auctions is an essential task for modern advertising auction systems. The main approach to managing digital auctions is an autobidding approach, which depends on the Click-Through Rate and Conversion Rate…

计算机科学与博弈论 · 计算机科学 2025-10-13 Andrey Pudovikov , Alexandra Khirianova , Ekaterina Solodneva , Gleb Molodtsov , Aleksandr Katrutsa , Yuriy Dorn , Egor Samosvat

We study an online learning problem on dynamic pricing and resource allocation, where we make joint pricing and inventory decisions to maximize the overall net profit. We consider the stochastic dependence of demands on the price, which…

机器学习 · 计算机科学 2025-05-23 Jianyu Xu , Xuan Wang , Yu-Xiang Wang , Jiashuo Jiang

In display advertising, a small group of sellers and bidders face each other in up to 10 12 auctions a day. In this context, revenue maximisation via monopoly price learning is a high-value problem for sellers. By nature, these auctions are…

机器学习 · 计算机科学 2020-10-21 Lorenzo Croissant , Marc Abeille , Clément Calauzènes

We study the problem of learning a linear model to set the reserve price in an auction, given contextual information, in order to maximize expected revenue from the seller side. First, we show that it is not possible to solve this problem…

最优化与控制 · 数学 2020-11-17 Joey Huchette , Haihao Lu , Hossein Esfandiari , Vahab Mirrokni

This paper studies the continuous-time reinforcement learning (RL) for optimal switching problems across multiple regimes. We consider a type of exploratory formulation under entropy regularization where the agent randomizes both the timing…

最优化与控制 · 数学 2025-12-23 Yijie Huang , Mengge Li , Xiang Yu , Zhou Zhou

Reinforcement learning-based recommender systems have recently gained popularity. However, due to the typical limitations of simulation environments (e.g., data inefficiency), most of the work cannot be broadly applied in all domains. To…

信息检索 · 计算机科学 2024-06-04 Xiaocong Chen , Siyu Wang , Lina Yao

This study investigates the development of an optimal execution strategy through reinforcement learning, aiming to determine the most effective approach for traders to buy and sell inventory within a finite time horizon. Our proposed model…

交易与市场微观结构 · 定量金融 2025-11-04 Yadh Hafsi , Edoardo Vittori

Fairness plays a crucial role in various multi-agent systems (e.g., communication networks, financial markets, etc.). Many multi-agent dynamical interactions can be cast as Markov Decision Processes (MDPs). While existing research has…

机器学习 · 计算机科学 2023-06-02 Peizhong Ju , Arnob Ghosh , Ness B. Shroff

We study the problem of finding the optimal bidding strategy for an advertiser in a multi-platform auction setting. The competition on a platform is captured by a value and a cost function, mapping bidding strategies to value and cost…

计算机科学与博弈论 · 计算机科学 2025-02-27 Gagan Aggarwal , Anupam Gupta , Xizhi Tan , Mingfei Zhao

Most e-commerce product feeds provide blended results of advertised products and recommended products to consumers. The underlying advertising and recommendation platforms share similar if not exactly the same set of candidate products.…

机器学习 · 统计学 2019-08-20 Dagui Chen , Junqi Jin , Weinan Zhang , Fei Pan , Lvyin Niu , Chuan Yu , Jun Wang , Han Li , Jian Xu , Kun Gai

Various methods for solving the inverse reinforcement learning (IRL) problem have been developed independently in machine learning and economics. In particular, the method of Maximum Causal Entropy IRL is based on the perspective of entropy…

机器学习 · 计算机科学 2021-03-05 Navyata Sanghvi , Shinnosuke Usami , Mohit Sharma , Joachim Groeger , Kris Kitani

In this paper, we propose a max-min entropy framework for reinforcement learning (RL) to overcome the limitation of the soft actor-critic (SAC) algorithm implementing the maximum entropy RL in model-free sample-based learning. Whereas the…

机器学习 · 计算机科学 2021-12-21 Seungyul Han , Youngchul Sung