English
Related papers

Related papers: Short-Long Policy Evaluation with Novel Actions

200 papers

Self-assessment rules play an essential role in safe and effective real-world robotic applications, which verify the feasibility of the selected action before actual execution. But how to utilize the self-assessment results to re-choose…

Robotics · Computer Science 2023-02-28 Kechun Xu , Runjian Chen , Shuqi Zhao , Zizhang Li , Hongxiang Yu , Ci Chen , Yue Wang , Rong Xiong

Following the rapid increase in Artificial Intelligence (AI) capabilities in recent years, the AI community has voiced concerns regarding possible safety risks. To support decision-making on the safe use and development of AI systems, there…

Machine Learning · Computer Science 2025-04-01 Gil Gekker , Meirav Segal , Dan Lahav , Omer Nevo

How do product teams evaluate LLM-powered products? As organizations integrate large language models (LLMs) into digital products, their unpredictable nature makes traditional evaluation approaches inadequate, yet little is known about how…

Software Engineering · Computer Science 2026-04-21 Willem van der Maden , Malak Sadek , Ziang Xiao , Aske Mottelson , Q. Vera Liao , Jichen Zhu

This paper focuses on developing Pareto-optimal estimation and policy learning to identify the most effective treatment that maximizes the total reward from both short-term and long-term effects, which might conflict with each other. For…

Machine Learning · Computer Science 2024-03-13 Yingrong Wang , Anpeng Wu , Haoxuan Li , Weiming Liu , Qiaowei Miao , Ruoxuan Xiong , Fei Wu , Kun Kuang

We explore the promises and challenges of employing sequential decision-making algorithms -- such as bandits, reinforcement learning, and active learning -- in law and public policy. While such algorithms have well-characterized performance…

Computers and Society · Computer Science 2022-11-30 Peter Henderson , Ben Chugg , Brandon Anderson , Daniel E. Ho

High-stakes decision domains are increasingly exploring the potential of Large Language Models (LLMs) for complex decision-making tasks. However, LLM deployment in real-world settings presents challenges in data security, evaluation of its…

Computers and Society · Computer Science 2025-12-05 Swati Sachan , Theo Miller , Mai Phuong Nguyen

A common dilemma encountered by many upon implementing an optimization method or experiment, whether it be a reinforcement learning algorithm, or A/B testing, is deciding on what metric to optimize for. Very often short-term metrics, which…

Applications · Statistics 2019-06-17 Yoni Schamroth , Liron Gat Kahlon , Boris Rabinovich , David Steinberg

We present relay policy learning, a method for imitation and reinforcement learning that can solve multi-stage, long-horizon robotic tasks. This general and universally-applicable, two-phase approach consists of an imitation learning stage…

Machine Learning · Computer Science 2019-10-29 Abhishek Gupta , Vikash Kumar , Corey Lynch , Sergey Levine , Karol Hausman

Applications of machine learning (ML) to high-stakes policy settings -- such as education, criminal justice, healthcare, and social service delivery -- have grown rapidly in recent years, sparking important conversations about how to ensure…

Machine Learning · Computer Science 2021-05-14 Hemank Lamba , Kit T. Rodolfa , Rayid Ghani

Most existing notions of algorithmic fairness are one-shot: they ensure some form of allocative equality at the time of decision making, but do not account for the adverse impact of the algorithmic decisions today on the long-term welfare…

Computers and Society · Computer Science 2019-06-28 Hoda Heidari , Vedant Nanda , Krishna P. Gummadi

In business and marketing, analyzing the reasons behind buying is a fundamental step towards understanding consumer behaviors, shaping business strategies, and predicting market outcomes. Prior research on purchase reason has relied on…

Information Retrieval · Computer Science 2024-11-19 Tao Chen , Siqi Zuo , Cheng Li , Mingyang Zhang , Qiaozhu Mei , Michael Bendersky

Off-policy policy evaluation methods for sequential decision making can be used to help identify if a proposed decision policy is better than a current baseline policy. However, a new decision policy may be better than a baseline policy for…

Machine Learning · Computer Science 2021-11-30 Ramtin Keramati , Omer Gottesman , Leo Anthony Celi , Finale Doshi-Velez , Emma Brunskill

Due to the rapidly rising popularity of Massive Open Online Courses (MOOCs), there is a growing demand for scalable automated support technologies for student learning. Transferring traditional educational resources to online contexts has…

Human-Computer Interaction · Computer Science 2018-09-13 Yohan Jo , Keith Maki , Gaurav Tomar

Longitudinal data often contains outcomes measured at multiple visits and scientific interest may lie in quantifying the effect of an intervention on an outcome's rate of change. For example, one may wish to study the progression (or…

Methodology · Statistics 2025-10-07 Anja Shahu , Weijie Xia , Ying Wei , Daniel Malinsky

We address the personalized policy learning problem using longitudinal mobile health application usage data. Personalized policy represents a paradigm shift from developing a single policy that may prescribe personalized decisions by…

Methodology · Statistics 2020-01-13 Xinyu Hu , Min Qian , Bin Cheng , Ying Kuen Cheung

Current human-AI alignment and evaluation methods for large language models (LLMs) often rely on preference signals collected immediately after an interaction. This practice implicitly treats preference as static, even though many…

Human-Computer Interaction · Computer Science 2026-05-06 Simret Araya Gebreegziabher , Allison E Sproul , Yinuo Yang , Chaoran Chen , Diego Gómez-Zará , Toby Jia-Jun Li

The goal of this article is to investigate how human participants allocate their limited time to decisions with different properties. We report the results of two behavioral experiments. In each trial of the experiments, the participant…

Neurons and Cognition · Quantitative Biology 2016-07-20 Arash Khodadadi , Pegah Fakhari , Jerome R. Busemeyer

Safety and responsibility evaluations of advanced AI models are a critical but developing field of research and practice. In the development of Google DeepMind's advanced AI models, we innovated on and applied a broad set of approaches to…

Policy evaluation estimates the performance of a policy by (1) collecting data from the environment and (2) processing raw data into a meaningful estimate. Due to the sequential nature of reinforcement learning, any improper data-collecting…

Machine Learning · Computer Science 2025-03-21 Shuze Daniel Liu , Claire Chen , Shangtong Zhang

Measuring innovation often relies on context-specific proxies and on expert evaluation. Hence, empirical innovation research is often limited to settings where such data is available. We investigate how large language models (LLMs) can be…

Computation and Language · Computer Science 2025-08-05 Robin Nowak , Patrick Figge , Carolin Haeussler
‹ Prev 1 4 5 6 7 8 10 Next ›