中文
相关论文

相关论文: Improving User Experience in Preference-Based Opti…

200 篇论文

In sequential recommendation, models recommend items based on user's interaction history. To this end, current models usually incorporate information such as item descriptions and user intent or preferences. User preferences are usually not…

When operating in service of people, robots need to optimize rewards aligned with end-user preferences. Since robots will rely on raw perceptual inputs like RGB images, their rewards will inevitably use visual representations. Recently…

机器人学 · 计算机科学 2024-01-17 Ran Tian , Chenfeng Xu , Masayoshi Tomizuka , Jitendra Malik , Andrea Bajcsy

Shared autonomy holds promise for improving the usability and accessibility of assistive robotic arms, but current methods often rely on costly expert demonstrations and remain static after pretraining, limiting their ability to handle…

机器人学 · 计算机科学 2025-07-28 Yiran Tao , Guixiu Qiao , Dan Ding , Zackory Erickson

A growing field in robotics and Artificial Intelligence (AI) research is human-robot collaboration, whose target is to enable effective teamwork between humans and robots. However, in many situations human teams are still superior to…

机器人学 · 计算机科学 2017-11-27 Giovanni Saponaro , Lorenzo Jamone , Alexandre Bernardino , Giampiero Salvi

Preference-based reinforcement learning (PbRL) is a suitable approach for style adaptation of pre-trained robotic behavior: adapting the robot's policy to follow human user preferences while still being able to perform the original task.…

During the preference optimization of large language models (LLMs), distribution shifts may arise between newly generated model samples and the data used to train the reward model (RM). This shift reduces the efficacy of the RM, which in…

机器学习 · 计算机科学 2025-06-11 Tianyuan Shi , Canbin Huang , Fanqi Wan , Longguang Zhong , Ziyi Yang , Weizhou Shen , Xiaojun Quan , Ming Yan

Flexible manufacturing processes demand robots to easily adapt to changes in the environment and interact with humans. In such dynamic scenarios, robotic tasks may be programmed through learning-from-demonstration approaches, where a…

机器人学 · 计算机科学 2019-08-21 Leonel Rozo

Recommendations are commonly used to modify user's natural behavior, for example, increasing product sales or the time spent on a website. This results in a gap between the ultimate business objective and the classical setup where…

信息检索 · 计算机科学 2019-05-23 Stephen Bonner , Flavian Vasile

Robots need models of human behavior for both inferring human goals and preferences, and predicting what people will do. A common model is the Boltzmann noisily-rational decision model, which assumes people approximately optimize a reward…

机器人学 · 计算机科学 2020-01-14 Andreea Bobu , Dexter R. R. Scobee , Jaime F. Fisac , S. Shankar Sastry , Anca D. Dragan

Human-robot handovers are characterized by high uncertainty and poor structure of the problem that make them difficult tasks. While machine learning methods have shown promising results, their application to problems with large state…

机器人学 · 计算机科学 2016-10-18 Francesco Riccio , Roberto Capobianco , Daniele Nardi

Human-robot interactions (HRI) can be modeled as dynamic or differential games with incomplete information, where each agent holds private reward parameters. Due to the open challenge in finding perfect Bayesian equilibria of such games,…

机器人学 · 计算机科学 2020-11-05 Yi Chen , Lei Zhang , Tanner Merry , Sunny Amatya , Wenlong Zhang , Yi Ren

In this paper, we propose a theoretically founded sequential strategy for training large-scale Recommender Systems (RS) over implicit feedback, mainly in the form of clicks. The proposed approach consists in minimizing pairwise ranking loss…

Modern information retrieval systems, including web search, ads placement, and recommender systems, typically rely on learning from user feedback. Click models, which study how users interact with a ranked list of items, provide a useful…

信息检索 · 计算机科学 2021-04-20 Xinyi Dai , Jianghao Lin , Weinan Zhang , Shuai Li , Weiwen Liu , Ruiming Tang , Xiuqiang He , Jianye Hao , Jun Wang , Yong Yu

Our goal in this paper is to plan the motion of a robot in a partitioned environment with dynamically changing, locally sensed rewards. We assume that arbitrary assumptions on the reward dynamics can be given. The robot aims to accomplish a…

机器人学 · 计算机科学 2012-08-30 Maria Svorenova , Jana Tumova , Jiri Barnat , Ivana Cerna

Preference-based Reinforcement Learning (PbRL) enables policy learning through simple queries comparing trajectories from a single policy. While human responses to these queries make it possible to learn policies aligned with human…

机器人学 · 计算机科学 2026-01-22 Yuki Kadokawa , Jonas Frey , Takahiro Miki , Takamitsu Matsubara , Marco Hutter

Parameter adaptation, that is the capability to automatically adjust an algorithm's hyperparameters depending on the problem being faced, is one of the main trends in evolutionary computation applied to numerical optimization. While several…

神经与进化计算 · 计算机科学 2022-06-30 Michele Tessari , Giovanni Iacca

Quadruped robots are showing impressive abilities to navigate the real world. If they are to become more integrated into society, social trust in interactions with humans will become increasingly important. Additionally, robots will need to…

机器人学 · 计算机科学 2024-07-01 Alessandra Chappuis , Guillaume Bellegarda , Auke Ijspeert

Machine learning models are being increasingly deployed to take, or assist in taking, complicated and high-impact decisions, from quasi-autonomous vehicles to clinical decision support systems. This poses challenges, particularly when…

机器学习 · 计算机科学 2023-11-14 Alex J. Chan , Alihan Huyuk , Mihaela van der Schaar

Herein we suggest a mobile robot-training algorithm that is based on the preference approximation of the decision taker who controls the robot, which in its turn is managed by the Markov chain. Setup of the model parameters is made on the…

机器人学 · 计算机科学 2015-09-07 Valery Vilisov

Reward modeling is central to alignment pipelines such as RLHF, RLAIF, and PPO-based policy optimization, yet its reliability is constrained by limited and heterogeneous human preference data that are expensive to collect at scale. While…

机器学习 · 计算机科学 2026-05-26 Payel Bhattacharjee , Osvaldo Simeone , Ravi Tandon