中文
相关论文

相关论文: Preference-based Teaching

200 篇论文

Preference-based reinforcement learning (PbRL) is an approach that enables RL agents to learn from preference, which is particularly useful when formulating a reward function is challenging. Existing PbRL methods generally involve a…

机器学习 · 计算机科学 2023-10-30 Gaon An , Junhyeok Lee , Xingdong Zuo , Norio Kosaka , Kyung-Min Kim , Hyun Oh Song

We present the model theoretic concepts that allow mathematics to be developed with the notion of the potential infinite instead of the actual infinite. The potential infinite is understood as a dynamic notion, being an indefinitely…

逻辑 · 数学 2022-12-16 Matthias Eberl

Large language models (LLMs) are susceptible to persuasion, which can pose risks when models are faced with an adversarial interlocutor. We take a first step towards defending models against persuasion while also arguing that defense…

计算与语言 · 计算机科学 2025-02-11 Elias Stengel-Eskin , Peter Hase , Mohit Bansal

We study a model of machine teaching where the teacher mapping is constructed from a size function on both concepts and examples. The main question in machine teaching is the minimum number of examples needed for any concept, the so-called…

组合数学 · 数学 2024-02-12 Brigt Håvardstun , Jan Kratochvíl , Joakim Sunde , Jan Arne Telle

Quite recently a teaching model, called "No-Clash Teaching" or simply "NC-Teaching", had been suggested that is provably optimal in the following strong sense. First, it satisfies Goldman and Matthias' collusion-freeness condition. Second,…

组合数学 · 数学 2022-05-06 Hans U. Simon

We consider the problem of learning preferences over trajectories for mobile manipulators such as personal robots and assembly line robots. The preferences we learn are more intricate than simple geometric constraints on trajectories; they…

机器人学 · 计算机科学 2016-01-06 Ashesh Jain , Shikhar Sharma , Thorsten Joachims , Ashutosh Saxena

We investigate the Plackett-Luce (PL) model based listwise learning-to-rank (LTR) on data with partitioned preference, where a set of items are sliced into ordered and disjoint partitions, but the ranking of items within a partition is…

机器学习 · 计算机科学 2021-03-01 Jiaqi Ma , Xinyang Yi , Weijing Tang , Zhe Zhao , Lichan Hong , Ed H. Chi , Qiaozhu Mei

Preference-based reinforcement learning (PbRL) shows promise in aligning robot behaviors with human preferences, but its success depends heavily on the accurate modeling of human preferences through reward models. Most methods adopt…

机器人学 · 计算机科学 2025-03-12 Dezhong Zhao , Ruiqi Wang , Dayoon Suh , Taehyeon Kim , Ziqin Yuan , Byung-Cheol Min , Guohua Chen

Choice functions accept a set of alternatives as input and produce a preferred subset of these alternatives as output. We study the problem of learning such functions under conditions of context-dependence of preferences, which means that…

机器学习 · 计算机科学 2021-10-25 Karlson Pfannschmidt , Pritha Gupta , Björn Haddenhorst , Eyke Hüllermeier

Machine teaching is an algorithmic framework for teaching a target hypothesis via a sequence of examples or demonstrations. We investigate machine teaching for temporal logic formulas -- a novel and expressive hypothesis class amenable to…

人工智能 · 计算机科学 2020-01-28 Zhe Xu , Yuxin Chen , Ufuk Topcu

Probably Approximately Correct (i.e., PAC) learning is a core concept of sample complexity theory, and efficient PAC learnability is often seen as a natural counterpart to the class P in classical computational complexity. But while the…

计算复杂性 · 计算机科学 2023-04-28 Cornelius Brand , Robert Ganian , Kirill Simonov

This paper investigates simultaneous preference and metric learning from a crowd of respondents. A set of items represented by $d$-dimensional feature vectors and paired comparisons of the form ``item $i$ is preferable to item $j$'' made by…

机器学习 · 统计学 2022-07-11 Gregory Canal , Blake Mason , Ramya Korlakai Vinayak , Robert Nowak

In this paper we consider multi-objective reinforcement learning where the objectives are balanced using preferences. In practice, the preferences are often given in an adversarial manner, e.g., customers can be picky in many applications.…

机器学习 · 计算机科学 2021-10-29 Jingfeng Wu , Vladimir Braverman , Lin F. Yang

Deep neural networks often contain far more parameters than training examples, yet they still manage to generalize well in practice. Classical complexity measures such as VC-dimension or PAC-Bayes bounds usually become vacuous in this…

机器学习 · 计算机科学 2025-08-26 Aviral Dhingra

We consider a school choice matching model where the priorities for schools are represented by binary relations that may not be weak order. We focus on the (total order) extensions of the binary relations. We introduce a class of algorithms…

理论经济学 · 经济学 2023-10-13 Minoru Kitahara , Yasunori Okumura

We consider the problem of learning to optimize an unknown Markov decision process (MDP). We show that, if the MDP can be parameterized within some known function class, we can obtain regret bounds that scale with the dimensionality, rather…

机器学习 · 统计学 2014-11-04 Ian Osband , Benjamin Van Roy

Transductive learning considers situations when a learner observes $m$ labelled training points and $u$ unlabelled test points with the final goal of giving correct answers for the test points. This paper introduces a new complexity measure…

机器学习 · 统计学 2016-02-24 Ilya Tolstikhin , Nikita Zhivotovskiy , Gilles Blanchard

Preference learning in Large Language Models (LLMs) has advanced significantly, yet existing methods remain limited by modest performance gains, high computational costs, hyperparameter sensitivity, and insufficient modeling of global…

计算与语言 · 计算机科学 2026-04-03 Liang Zhu , Yuelin Bai , Xiankun Ren , Jiaxi Yang , Lei Zhang , Feiteng Fang , Hamid Alinejad-Rokny , Minghuan Tan , Min Yang

In practice, preference learning from human feedback depends on incomplete data with hidden context. Hidden context refers to data that affects the feedback received, but which is not represented in the data used to train a preference…

机器学习 · 计算机科学 2024-04-18 Anand Siththaranjan , Cassidy Laidlaw , Dylan Hadfield-Menell

Recent advances in large language models (LLMs) have significantly improved the alignment of models with general human preferences. However, a major challenge remains in adapting LLMs to individual preferences, which are not only diverse…

计算与语言 · 计算机科学 2026-04-15 Shanyong Wang , Shuhang Lin , Yining Zhao , Xi Zhu , Yongfeng Zhang