中文
相关论文

相关论文: Risk-averse Batch Active Inverse Reward Design

200 篇论文

Response adaptive randomization (RAR) is appealing from methodological, ethical, and pragmatic perspectives in the sense that subjects are more likely to be randomized to better performing treatment groups based on accumulating data.…

统计方法学 · 统计学 2022-08-03 Tianyu Zhan , Lu Cui , Ziqian Geng , Lanju Zhang , Yihua Gu , Ivan S. F. Chan

This paper proposes risk-averse and risk-agnostic formulations to robust design in which solutions that satisfy the system requirements for a set of scenarios are pursued. These scenarios, which correspond to realizations of uncertain…

最优化与控制 · 数学 2025-11-07 Luis G. Crespo , Bret Stanford , Natalia Alexandrov

In high-stakes AI applications, even a single action can cause irreparable damage. However, nearly all of sequential decision-making theory assumes that all errors are recoverable (e.g., by bounding rewards). Standard bandit algorithms that…

机器学习 · 计算机科学 2026-04-14 Sarah Liaw , Benjamin Plaut

We consider nonstationary multi-armed bandit problems where the model parameters of the arms change over time. We introduce the adaptive resetting bandit (ADR-bandit), a bandit algorithm class that leverages adaptive windowing techniques…

机器学习 · 统计学 2023-10-27 Junpei Komiyama , Edouard Fouché , Junya Honda

The ability to accurately predict the trajectory of surrounding vehicles is a critical hurdle to overcome on the journey to fully autonomous vehicles. To address this challenge, we pioneer a novel behavior-aware trajectory prediction model…

机器人学 · 计算机科学 2023-12-18 Haicheng Liao , Zhenning Li , Huanming Shen , Wenxuan Zeng , Dongping Liao , Guofa Li , Shengbo Eben Li , Chengzhong Xu

Data generation and labeling are often expensive in robot learning. Preference-based learning is a concept that enables reliable labeling by querying users with preference questions. Active querying methods are commonly employed in…

机器学习 · 计算机科学 2024-02-27 Erdem Bıyık , Nima Anari , Dorsa Sadigh

A well-defined reward function is crucial for successful training of an reinforcement learning (RL) agent. However, defining a suitable reward function is a notoriously challenging task, especially in complex, multi-objective environments.…

人工智能 · 计算机科学 2023-08-31 Jasmina Gajcin , James McCarthy , Rahul Nair , Radu Marinescu , Elizabeth Daly , Ivana Dusparic

Deducing the contribution of each agent and assigning the corresponding reward to them is a crucial problem in cooperative Multi-Agent Reinforcement Learning (MARL). Previous studies try to resolve the issue through designing an intrinsic…

机器学习 · 计算机科学 2023-02-21 Wei Li , Weiyan Liu , Shitong Shao , Shiyi Huang

Reward learning enables robots to learn adaptable behaviors from human input. Traditional methods model the reward as a linear function of hand-crafted features, but that requires specifying all the relevant features a priori, which is…

机器人学 · 计算机科学 2022-01-19 Andreea Bobu , Marius Wiggert , Claire Tomlin , Anca D. Dragan

Robot navigation in dynamic, crowded environments poses a significant challenge due to the inherent uncertainties in the obstacle model. In this work, we propose a risk-adaptive approach based on the Conditional Value-at-Risk Barrier…

机器人学 · 计算机科学 2025-08-04 Xinyi Wang , Taekyung Kim , Bardh Hoxha , Georgios Fainekos , Dimitra Panagou

Classical multi-armed bandit problems use the expected value of an arm as a metric to evaluate its goodness. However, the expected value is a risk-neutral metric. In many applications like finance, one is interested in balancing the…

机器学习 · 计算机科学 2019-06-04 Anmol Kagrecha , Jayakrishnan Nair , Krishna Jagannathan

Our ultimate goal is to build robust policies for robots that assist people. What makes this hard is that people can behave unexpectedly at test time, potentially interacting with the robot outside its training distribution and leading to…

人工智能 · 计算机科学 2023-10-17 Jerry Zhi-Yang He , Zackory Erickson , Daniel S. Brown , Anca D. Dragan

We study the reward-free reinforcement learning framework, which is particularly suitable for batch reinforcement learning and scenarios where one needs policies for multiple reward functions. This framework has two phases. In the…

机器学习 · 计算机科学 2020-10-26 Zihan Zhang , Simon S. Du , Xiangyang Ji

We seek to align agent policy with human expert behavior in a reinforcement learning (RL) setting, without any prior knowledge about dynamics, reward function, and unsafe states. There is a human expert knowing the rewards and unsafe states…

机器学习 · 计算机科学 2020-01-01 Daniel Hsu

We introduce Deep Adaptive Design (DAD), a method for amortizing the cost of adaptive Bayesian experimental design that allows experiments to be run in real-time. Traditional sequential Bayesian optimal experimental design approaches…

机器学习 · 统计学 2021-06-14 Adam Foster , Desi R. Ivanova , Ilyas Malik , Tom Rainforth

Restless multi-armed bandits (RMAB) play a central role in modeling sequential decision making problems under an instantaneous activation constraint that at most B arms can be activated at any decision epoch. Each restless arm is endowed…

机器学习 · 计算机科学 2024-05-03 Guojun Xiong , Jian Li

In human-robot interaction (HRI) systems, such as autonomous vehicles, understanding and representing human behavior are important. Human behavior is naturally rich and diverse. Cost/reward learning, as an efficient way to learn and…

机器人学 · 计算机科学 2020-08-24 Liting Sun , Zheng Wu , Hengbo Ma , Masayoshi Tomizuka

We consider the problem of reward learning for temporally extended tasks. For reward learning, inverse reinforcement learning (IRL) is a widely used paradigm. Given a Markov decision process (MDP) and a set of demonstrations for a task, IRL…

机器人学 · 计算机科学 2021-07-14 Farzan Memarian , Zhe Xu , Bo Wu , Min Wen , Ufuk Topcu

We consider designing reward schemes that incentivize agents to create high-quality content (e.g., videos, images, text, ideas). The problem is at the center of a real-world application where the goal is to optimize the overall quality of…

计算机科学与博弈论 · 计算机科学 2022-05-03 Mengjing Chen , Pingzhong Tang , Zihe Wang , Shenke Xiao , Xiwang Yang

Adversarial attacks in 3D environments have emerged as a critical threat to the reliability of visual perception systems, particularly in safety-sensitive applications such as identity verification and autonomous driving. These attacks…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Xiao Yang , Lingxuan Wu , Lizhong Wang , Chengyang Ying , Hang Su , Jun Zhu