中文
相关论文

相关论文: Optimizing an Utility Function for Exploration / E…

200 篇论文

Online learning to rank (OL2R) has attracted great research interests in recent years, thanks to its advantages in avoiding expensive relevance labeling as required in offline supervised ranking model learning. Such a solution explores the…

信息检索 · 计算机科学 2021-11-02 Yiling Jia , Hongning Wang

This paper examines the exploration-exploitation trade-off in reinforcement learning with verifiable rewards (RLVR), a framework for improving the reasoning of Large Language Models (LLMs). Recent studies suggest that RLVR can elicit strong…

机器学习 · 计算机科学 2026-01-27 Peter Chen , Xiaopeng Li , Ziniu Li , Wotao Yin , Xi Chen , Tianyi Lin

This study investigates the development of an optimal execution strategy through reinforcement learning, aiming to determine the most effective approach for traders to buy and sell inventory within a finite time horizon. Our proposed model…

交易与市场微观结构 · 定量金融 2025-11-04 Yadh Hafsi , Edoardo Vittori

Reinforcement learning can greatly benefit from the use of options as a way of encoding recurring behaviours and to foster exploration. An important open problem is how can an agent autonomously learn useful options when solving particular…

机器学习 · 计算机科学 2020-01-07 Manuel Del Verme , Bruno Castro da Silva , Gianluca Baldassarre

Exploration strategy design is one of the challenging problems in reinforcement learning~(RL), especially when the environment contains a large state space or sparse rewards. During exploration, the agent tries to discover novel areas or…

机器学习 · 计算机科学 2019-06-07 Xiao Ma , Shen-Yi Zhao , Wu-Jun Li

Decentralized resource allocation is a key problem for large-scale autonomic (or self-managing) computing systems. Motivated by a data center scenario, we explore efficient techniques for resolving resource conflicts via cooperative…

计算机科学与博弈论 · 计算机科学 2012-12-12 Craig Boutilier , Rajarshi Das , Jeffrey O. Kephart , Gerald Tesauro , William E. Walsh

Many real-world problems contain multiple objectives and agents, where a trade-off exists between objectives. Key to solving such problems is to exploit sparse dependency structures that exist between agents. For example, in wind farm…

人工智能 · 计算机科学 2022-07-04 Conor F. Hayes , Timothy Verstraeten , Diederik M. Roijers , Enda Howley , Patrick Mannion

We examine trade-offs among stakeholders in ad auctions. Our metrics are the revenue for the utility of the auctioneer, the number of clicks for the utility of the users and the welfare for the utility of the advertisers. We show how to…

计算机科学与博弈论 · 计算机科学 2014-04-22 Yoram Bachrach , Sofia Ceppi , Ian A. Kash , Peter Key , David Kurokawa

In many real-world applications of reinforcement learning (RL), performing actions requires consuming certain types of resources that are non-replenishable in each episode. Typical applications include robotic control with limited energy…

机器学习 · 计算机科学 2022-12-15 Zhihai Wang , Taoxing Pan , Qi Zhou , Jie Wang

We consider reinforcement learning (RL) in continuous time and study the problem of achieving the best trade-off between exploration of a black box environment and exploitation of current knowledge. We propose an entropy-regularized reward…

最优化与控制 · 数学 2019-02-14 Haoran Wang , Thaleia Zariphopoulou , Xunyu Zhou

Conversational Recommender Systems (CRSs) aim to provide personalized recommendations through multi-turn natural language interactions with users. Given the strong interaction and reasoning skills of Large Language Models (LLMs), leveraging…

计算与语言 · 计算机科学 2025-10-02 Xiaoyan Zhao , Ming Yan , Yang Zhang , Yang Deng , Jian Wang , Fengbin Zhu , Yilun Qiu , Hong Cheng , Tat-Seng Chua

How do people navigate the exploration-exploitation (EE) trade-off when making repeated choices with unknown rewards? We study this question through the lens of multi-armed bandit problems and introduce a novel behavioral model, Quantal…

最优化与控制 · 数学 2024-12-25 Jingying Ding , Yifan Feng , Ying Rong

We present a deterministic exploration mechanism for sponsored search auctions, which enables the auctioneer to learn the relevance scores of advertisers, and allows advertisers to estimate the true value of clicks generated at the auction…

计算机科学与博弈论 · 计算机科学 2011-11-10 Sudhir Kumar Singh , Vwani P. Roychowdhury , Milan Bradonjić , Behnam A. Rezaei

In networked environments, users frequently share recommendations about content, products, services, and courses of action with others. The extent to which such recommendations are successful and adopted is highly contextual, dependent on…

机器学习 · 计算机科学 2025-10-23 Ahmed Sayeed Faruk , Mohammad Shahverdikondori , Elena Zheleva

Personalized news recommendation aims to provide attractive articles for readers by predicting their likelihood of clicking on a certain article. To accurately predict this probability, plenty of studies have been proposed that actively…

信息检索 · 计算机科学 2021-12-30 Sungmin Cho , Hongjun Lim , Keunchan Park , Sungjoo Yoo , Eunhyeok Park

Recent advancements in agentic test-time scaling allow models to gather environmental feedback before committing to final actions. A key limitation of existing methods is that they typically employ undifferentiated exploration strategies,…

人工智能 · 计算机科学 2026-05-13 Xingyuan Hua , Sheng Yue , Ju Ren

How do you incentivize self-interested agents to $\textit{explore}$ when they prefer to $\textit{exploit}$? We consider complex exploration problems, where each agent faces the same (but unknown) MDP. In contrast with traditional…

机器学习 · 计算机科学 2023-02-21 Max Simchowitz , Aleksandrs Slivkins

Information foraging connects optimal foraging theory in ecology with how humans search for information. The theory suggests that, following an information scent, the information seeker must optimize the tradeoff between exploration by…

信息检索 · 计算机科学 2016-11-18 Peter Wittek , Ying-Hsang Liu , Sándor Darányi , Tom Gedeon , Ik Soo Lim

This work is about optimal order execution, where a large order is split into several small orders to maximize the implementation shortfall. Based on the diversity of cryptocurrency exchanges, we attempt to extract cross-exchange signals by…

交易与市场微观结构 · 定量金融 2023-07-03 Cong Zheng , Jiafa He , Can Yang

We investigate approximately optimal mechanisms in settings where bidders' utility functions are non-linear; specifically, convex, with respect to payments (such settings arise, for instance, in procurement auctions for energy). We provide…

计算机科学与博弈论 · 计算机科学 2017-02-23 Amy Greenwald , Takehiro Oyakawa , Vasilis Syrgkanis