中文
相关论文

相关论文: Provably Efficient Algorithm for Best Scoring Rule…

200 篇论文

Safe offline reinforcement learning aims to learn policies that maximize cumulative rewards while adhering to safety constraints, using only offline data for training. A key challenge is balancing safety and performance, particularly when…

机器学习 · 计算机科学 2024-12-13 Prajwal Koirala , Zhanhong Jiang , Soumik Sarkar , Cody Fleming

Age-of-Information (AoI) is a critical metric for network applications. Existing works mostly address optimization with homogeneous AoI requirements, which is different from practice. In this work, we optimize uplink scheduling for an…

系统与控制 · 电气工程与系统科学 2022-12-14 Shuang Wu , Xiaoqiang Ren , Qing-Shan Jia , Karl Henrik Johansson , Ling Shi

We propose a framework for adaptive data-centric collaborative machine learning among self-interested agents, coordinated by an arbiter. Designed to handle the incremental nature of real-world data, the framework operates in an online…

机器学习 · 计算机科学 2025-02-07 Nithia Vijayan , Bryan Kian Hsiang Low

Continually solving new, unsolved tasks is the key to learning diverse behaviors. Through reinforcement learning (RL), we have made massive strides towards solving tasks that have a single goal. However, in the multi-task domain, where an…

机器学习 · 计算机科学 2020-06-18 Yunzhi Zhang , Pieter Abbeel , Lerrel Pinto

The increasing deployment of AI is shaping the future landscape of the internet, which is set to become an integrated ecosystem of AI agents. Orchestrating the interaction among AI agents necessitates decentralized, self-sustaining…

计算机科学与博弈论 · 计算机科学 2024-10-08 Dima Ivanov , Paul Dütting , Inbal Talgam-Cohen , Tonghan Wang , David C. Parkes

We formulate the local ranking problem in the framework of bipartite ranking where the goal is to focus on the best instances. We propose a methodology based on the construction of real-valued scoring functions. We study empirical risk…

统计理论 · 数学 2016-08-16 Stéphan Clémençon , Nicolas Vayatis

Optimal power flow (OPF) is a very fundamental but vital optimization problem in the power system, which aims at solving a specific objective function (ex.: generator costs) while maintaining the system in the stable and safe operations. In…

系统与控制 · 电气工程与系统科学 2020-04-09 Yuhao Zhou , Bei Zhang , Chunlei Xu , Tu Lan , Ruisheng Diao , Di Shi , Zhiwei Wang , Wei-Jen Lee

We consider the principal-agent problem with heterogeneous agents. Previous works assume that the principal signs independent incentive contracts with every agent to make them invest more efforts on the tasks. However, in many…

多智能体系统 · 计算机科学 2019-11-12 Shenke Xiao , Zihe Wang , Mengjing Chen , Pingzhong Tang , Xiwang Yang

Online advertising has motivated interest in online selection problems. Displaying ads to the right users benefits both the platform (e.g., via pay-per-click) and the advertisers (by increasing their reach). In practice, not all users click…

计算机科学与博弈论 · 计算机科学 2024-08-16 Sebastian Perez-Salazar , Mohit Singh , Alejandro Toriello

In practice, incentive providers (i.e., principals) often cannot observe the reward realizations of incentivized agents, which is in contrast to many principal-agent models that have been previously studied. This information asymmetry…

机器学习 · 计算机科学 2023-08-15 Ilgin Dogan , Zuo-Jun Max Shen , Anil Aswani

We study the problems of offline and online contextual optimization with feedback information, where instead of observing the loss, we observe, after-the-fact, the optimal action an oracle with full knowledge of the objective function would…

机器学习 · 计算机科学 2023-07-04 Omar Besbes , Yuri Fonseca , Ilan Lobel

Augmenting the input of algorithms with predictions is an algorithm design paradigm that suggests leveraging a (possibly erroneous) prediction to improve worst-case performance guarantees when the prediction is perfect (consistency), while…

Our goal is to deploy a high-accuracy system starting with zero training examples. We consider an "on-the-job" setting, where as inputs arrive, we use real-time crowdsourcing to resolve uncertainty where needed and output our prediction…

人工智能 · 计算机科学 2015-12-09 Keenon Werling , Arun Chaganty , Percy Liang , Chris Manning

In recent years, machine learning has begun automating decision making in fields as varied as college admissions, credit lending, and criminal sentencing. The socially sensitive nature of some of these applications together with increasing…

机器学习 · 计算机科学 2021-07-06 Connor Lawless , Oktay Gunluk

We consider a principal agent project selection problem with asymmetric information. There are $N$ projects and the principal must select exactly one of them. Each project provides some profit to the principal and some payoff to the agent…

理论经济学 · 经济学 2025-04-15 Sumit Goel , Wade Hann-Caruthers

Offline reinforcement learning (RL) enables policy learning from static data but often suffers from poor coverage of the state-action space and distributional shift problems. This problem can be addressed by allowing limited online…

机器学习 · 计算机科学 2026-02-03 Soumyadeep Roy , Shashwat Kushwaha , Ambedkar Dukkipati

Recent developments in cyber-physical systems have increased the importance of maximizing the freshness of the information about the physical environment. However, optimizing the access policies of Internet of Things devices to maximize the…

网络与互联网体系结构 · 计算机科学 2025-12-16 José-Ramón Vidal , Vicent Pla , Luis Guijarro , Israel Leyva-Mayorga

We study an online version of the max-min fair allocation problem for indivisible items. In this problem, items arrive one by one, and each item must be allocated irrevocably on arrival to one of $n$ agents, who have additive valuations for…

计算机科学与博弈论 · 计算机科学 2021-11-16 Yasushi Kawase , Hanna Sumita

We study a setting in which a principal selects an agent to execute a collection of tasks according to a specified priority sequence. Agents, however, have their own individual priority sequences according to which they wish to execute the…

计算机科学与博弈论 · 计算机科学 2024-10-30 Donya G. Dobakhshari , Lav R. Varshney , Vijay Gupta

We investigate online scheduling with commitment for parallel identical machines. Our objective is to maximize the total processing time of accepted jobs. As soon as a job has been submitted, the commitment constraint forces us to decide…

数据结构与算法 · 计算机科学 2019-04-15 Chris Schwiegelshohn , Uwe Schwiegelshohn