中文
相关论文

相关论文: Optimal Anonymous Independent Reward Scheme Design

200 篇论文

In human-in-the-loop reinforcement learning or environments where calculating a reward is expensive, the costly rewards can make learning efficiency challenging to achieve. The cost of obtaining feedback from humans or calculating expensive…

机器学习 · 计算机科学 2025-03-03 Muhammed Yusuf Satici , David L. Roberts

Time-inconsistency refers to a paradox in decision making where agents exhibit inconsistent behaviors over time. Examples are procrastination where agents tends to costly postpone easy tasks, and abandonments where agents start a plan and…

计算机科学与博弈论 · 计算机科学 2015-04-03 Pingzhong Tang , Yifeng Teng , Zihe Wang , Shenke Xiao , Yichong Xu

We consider potentially non-convex optimization problems, for which optimal rates of approximation depend on the dimension of the parameter space and the smoothness of the function to be optimized. In this paper, we propose an algorithm…

机器学习 · 计算机科学 2022-04-12 Blake Woodworth , Francis Bach , Alessandro Rudi

The revenue optimal mechanism for selling a single item to agents with independent but non-identically distributed values is complex for agents with linear utility (Myerson,1981) and has no closed-form characterization for agents with…

计算机科学与博弈论 · 计算机科学 2020-11-24 Yiding Feng , Jason D. Hartline , Yingkai Li

This paper concerns the mechanism design for online resource allocation in a strategic setting. In this setting, a single supplier allocates capacity-limited resources to requests that arrive in a sequential and arbitrary manner. Each…

计算机科学与博弈论 · 计算机科学 2023-10-10 Xiaoqi Tan , Bo Sun , Alberto Leon-Garcia , Yuan Wu , Danny H. K. Tsang

We study the problem of fair online resource allocation via non-monetary mechanisms, where multiple agents repeatedly share a resource without monetary transfers. Previous work has shown that every agent can guarantee $1/2$ of their ideal…

计算机科学与博弈论 · 计算机科学 2025-05-27 David X. Lin , Daniel Hall , Giannis Fikioris , Siddhartha Banerjee , Éva Tardos

We consider multi-agent decision making, where each agent optimizes its cost function subject to constraints. Agents' actions belong to a compact convex Euclidean space and the agents' cost functions are coupled. We propose a distributed…

最优化与控制 · 数学 2016-12-01 Tatiana Tatarenko , Maryam Kamgarpour

In this paper, we study an unmanned aerial vehicle (UAV) communication system, where a ground node (GN) communicate with a UAV assisted by intelligent reflecting surface (IRS) in the presence of a jammer with imperfect location information.…

信息论 · 计算机科学 2022-01-25 Zhi Ji , Xinrong Guan , Jia Tu , Qingqing Wu , Wendong Yang

We study the design of effort-maximizing grading schemes between agents with private abilities. Assuming agents derive value from the information their grade reveals about their ability, we find that more informative grading schemes induce…

计算机科学与博弈论 · 计算机科学 2024-11-11 Sumit Goel

With the growing practical interest in vision-based tasks for autonomous systems, the need for efficient and complex methods becomes increasingly larger. In the rush to develop new methods with the aim to outperform the current state of the…

机器学习 · 计算机科学 2025-03-26 Daniel Yang

Buildings are a large consumer of energy, and reducing their energy usage may provide financial and societal benefits. One challenge in achieving efficient building operation is the fact that few financial motivations exist for encouraging…

最优化与控制 · 数学 2012-07-12 Anil Aswani , Claire Tomlin

Aligning Large Language Models (LLMs) to cater to different human preferences, learning new skills, and unlearning harmful behavior is an important problem. Search-based methods, such as Best-of-N or Monte-Carlo Tree Search, are performant,…

机器学习 · 计算机科学 2024-05-13 Seungwook Han , Idan Shenfeld , Akash Srivastava , Yoon Kim , Pulkit Agrawal

We identify two issues with the family of algorithms based on the Adversarial Imitation Learning framework. The first problem is implicit bias present in the reward functions used in these algorithms. While these biases might work well for…

机器学习 · 计算机科学 2018-10-16 Ilya Kostrikov , Kumar Krishna Agrawal , Debidatta Dwibedi , Sergey Levine , Jonathan Tompson

We investigate transmission optimization for intelligent reflecting surface (IRS) assisted multi-antenna systems from the physical-layer security perspective. The design goal is to maximize the system secrecy rate subject to the source…

信息论 · 计算机科学 2020-05-12 Hong Shen , Wei Xu , Shulei Gong , Zhenyao He , Chunming Zhao

To promote cooperation in Multi-Agent Reinforcement Learning, the reward signals of all agents can be aggregated together, forming global rewards that are commonly known as the fully cooperative setting. However, global rewards are usually…

机器学习 · 计算机科学 2026-01-30 Bang Giang Le , Viet Cuong Ta

Motivated by real-world applications such as the allocation of public housing, we examine the problem of assigning a group of agents to vertices (e.g., spatial locations) of a network so that the diversity level is maximized. Specifically,…

数据结构与算法 · 计算机科学 2024-04-02 Zirou Qiu , Andrew Yuan , Chen Chen , Madhav V. Marathe , S. S. Ravi , Daniel J. Rosenkrantz , Richard E. Stearns , Anil Vullikanti

Large language models (LLMs) have demonstrated remarkable capabilities across a range of text-generation tasks. However, LLMs still struggle with problems requiring multi-step decision-making and environmental feedback, such as online…

人工智能 · 计算机科学 2025-02-18 Zhenfang Chen , Delin Chen , Rui Sun , Wenjun Liu , Chuang Gan

AI systems often rely on two key components: a specified goal or reward function and an optimization algorithm to compute the optimal behavior for that goal. This approach is intended to provide value for a principal: the user on whose…

人工智能 · 计算机科学 2021-02-09 Simon Zhuang , Dylan Hadfield-Menell

In this paper we consider a distributed optimization scenario in which a set of agents has to solve a convex optimization problem with separable cost function, local constraint sets and a coupling inequality constraint. We propose a novel…

系统与控制 · 计算机科学 2018-04-25 Ivano Notarnicola , Giuseppe Notarstefano

We study the fair allocation of a cake, which serves as a metaphor for a divisible resource, under the requirement that each agent should receive a contiguous piece of the cake. While it is known that no finite envy-free algorithm exists in…

计算机科学与博弈论 · 计算机科学 2020-09-24 Paul W. Goldberg , Alexandros Hollender , Warut Suksompong