中文
相关论文

相关论文: How good are Popular Matchings?

200 篇论文

The stable marriage problem and its extensions have been extensively studied, with much of the work in the literature assuming that agents fully know their own preferences over alternatives. This assumption however is not always practical…

计算机科学与博弈论 · 计算机科学 2016-03-21 Baharak Rastegari , Paul Goldberg , David Manlove

Reinforcement Learning from Human Feedback (RLHF) is currently the leading approach for aligning large language models with human preferences. Typically, these models rely on extensive offline preference datasets for training. However,…

机器学习 · 计算机科学 2024-12-17 Avinandan Bose , Zhihan Xiong , Aadirupa Saha , Simon Shaolei Du , Maryam Fazel

As robots are being integrated into our daily lives, it becomes necessary to provide guarantees on the safe and provably correct operation. Such guarantees can be provided using automata theoretic task and mission planning where the…

系统与控制 · 计算机科学 2014-11-27 Kangjin Kim , Georgios E. Fainekos , Sriram Sankaranarayanan

The stable roommates problem with $n$ agents has worst case complexity $O(n^2)$ in time and space. Random instances can be solved faster and with less memory, however. We introduce an algorithm that has average time and space complexity…

数据结构与算法 · 计算机科学 2015-01-22 Stephan Mertens

When computing stable matchings, it is usually assumed that the preferences of the agents in the matching market are fixed. However, in many realistic scenarios, preferences change over time. Consequently, an initially stable matching may…

计算机科学与博弈论 · 计算机科学 2022-11-09 Niclas Boehmer , Klaus Heeger , Rolf Niedermeier

Standard RLHF relies on transitive scalar rewards, failing to capture the cyclic nature of human preferences. While some approaches like the General Preference Model (GPM) address this, we identify a theoretical limitation: their implicit…

计算与语言 · 计算机科学 2026-05-19 Yucong Huang , Xiucheng Li , Kaiqi Zhao , Jing Li

This paper investigates the relay selection (RS) problem in networks with multiple users and multiple common amplify-and-forward (AF) relays. Considering the overall quality-of-service of the network, we first specify our definition of…

信息论 · 计算机科学 2011-10-20 Saman Atapattu , Yindi Jing , Hai Jiang , Chintha Tellambura

A central issue lying at the heart of online reinforcement learning (RL) is data efficiency. While a number of recent works achieved asymptotically minimal regret in online RL, the optimality of these results is only guaranteed in a…

机器学习 · 计算机科学 2025-04-30 Zihan Zhang , Yuxin Chen , Jason D. Lee , Simon S. Du

This paper analyses the data rate achieved by various relay selection schemes in a single-user multi-hop relay network with decode-and-forward (DF) relaying. While the single-user relay selection problem is well studied in the literature,…

信息论 · 计算机科学 2023-08-17 Shalanika Dayarathna , Rajitha Senanayake , Jamie Evans

Justified representation (JR) and extended justified representation (EJR) are well-established proportionality axioms in approval-based multiwinner voting. Both axioms are always satisfiable, but they rely on a fixed quota (typically Hare…

计算机科学与博弈论 · 计算机科学 2026-02-18 Patrick Becker , Fabian Frank

We study the stable matching problem in non-bipartite graphs with incomplete but strict preference lists, where the edges have weights and the goal is to compute a stable matching of minimum or maximum weight. This problem is known to be…

计算机科学与博弈论 · 计算机科学 2017-03-28 Linda Farczadi , Natália Guričanová

Online and offline RLHF methods, such as PPO and DPO, have been highly successful in aligning AI with human preferences. Despite their success, however, these methods suffer from fundamental limitations: (a) Models trained with RLHF can…

机器学习 · 计算机科学 2025-04-15 Eugene Choi , Arash Ahmadian , Matthieu Geist , Oilvier Pietquin , Mohammad Gheshlaghi Azar

We investigate the maximum happy vertices (MHV) problem and its complement, the minimum unhappy vertices (MUHV) problem. We first show that the MHV and MUHV problems are a special case of the supermodular and submodular multi-labeling…

数据结构与算法 · 计算机科学 2017-01-12 Yao Xu , Peng Zhang , Randy Goebel , Guohui Lin

Reinforcement Learning from Human Feedback (RLHF) can reveal implicit objectives such as safety considerations that go beyond task completion. In this work, we focus on the common safety criteria embedded in crowd preference datasets, where…

人工智能 · 计算机科学 2026-05-22 Qian Lin , Daniel S. Brown

A Regret Minimizing Set (RMS) is a useful concept in which a smaller subset of a database is selected while mostly preserving the best scores along every possible utility function. In this paper, we study the $k$-Regret Minimizing Sets…

数据库 · 计算机科学 2022-01-19 Phoomraphee Luenam , Yau Pun Chen , Raymond Chi-Wing Wong

In this paper, we consider the multichannel rendezvous problem in cognitive radio networks (CRNs) where the probability that two users hopping on the same channel have a successful rendezvous is a function of channel states. The channel…

信息论 · 计算机科学 2019-06-26 Cheng-Shang Chang , Duan-Shin Lee , Yu-Lun Lin , Jen-Hung Wang

This paper has two main goals: (a) establish several statistical properties---consistency, asymptotic distributions, and convergence rates---of stationary solutions and values of a class of coupled nonconvex and nonsmoothempirical risk…

统计理论 · 数学 2019-10-08 Zhengling Qi , Ying Cui , Yufeng Liu , Jong-Shi Pang

Randomized Controlled Trials (RCTs) are the gold standard for comparing the effectiveness of a new treatment to the current one (the control). Most RCTs allocate the patients to the treatment group and the control group by uniform…

机器学习 · 统计学 2018-10-22 Onur Atan , William R. Zame , Mihaela van der Schaar

We study many-to-one matching problems between institutions and individuals, where each institution may be matched to multiple individuals. The matching market includes couples, who view pairs of institutions as complementary. Institutions'…

理论经济学 · 经济学 2025-07-11 Shashwat Khare , Souvik Roy

Adaptivity to changing environments and constraints is key to success in modern society. We address this by proposing "incrementalized versions" of Stable Marriage and Stable Roommates. That is, we try to answer the following question: for…

计算机科学与博弈论 · 计算机科学 2019-11-25 Robert Bredereck , Jiehua Chen , Dušan Knop , Junjie Luo , Rolf Niedermeier
‹ 上一页 1 8 9 10 下一页 ›