English
Related papers

Related papers: Risk Aware Benchmarking of Large Language Models

200 papers

As large language models (LLMs) expose systemic security challenges in high risk applications, including privacy leaks, bias amplification, and malicious abuse, there is an urgent need for a dynamic risk assessment and collaborative defence…

Cryptography and Security · Computer Science 2026-02-05 Xiaoyan Zhang , Dongyang Lyu , Xiaoqi Li

Considering the ever-evolving threat landscape and rapid changes in software development, we propose a risk assessment framework called SAFER (Software Analysis Framework for Evaluating Risk). This framework is based on the necessity of a…

Software Engineering · Computer Science 2024-12-25 Sarah Ali Siddiqui , Chandra Thapa , Rayne Holland , Wei Shao , Seyit Camtepe

In recent years, machine learning models have achieved great success at the expense of highly complex black-box structures. By using axiomatic attribution methods, we can fairly allocate the contributions of each feature, thus allowing us…

Computational Finance · Quantitative Finance 2025-06-10 Dangxing Chen

This paper introduces a rule for policy selection in the presence of estimation uncertainty, explicitly accounting for estimation risk. The rule belongs to the class of risk-aware rules on the efficient decision frontier, characterized as…

Econometrics · Economics 2026-01-21 Victor Chernozhukov , Sokbae Lee , Adam M. Rosen , Liyang Sun

We consider various stochastic models that incorporate the notion of risk-averseness into the standard 2-stage recourse model, and develop novel techniques for solving the algorithmic problems arising in these models. A key notable feature…

Data Structures and Algorithms · Computer Science 2008-05-06 Chaitanya Swamy

Model uncertainty has been one prominent issue both in the theory of risk measures and in practice such as financial risk management and regulation. Motivated by this observation, in this paper, we take a new perspective to describe the…

Theoretical Economics · Economics 2025-04-14 Shuo Gong , Yijun Hu , Linxiao Wei

Randomized controlled experiments assess new policy impacts on performance metrics to inform launch decisions. Traditional approaches evaluate metrics independently despite correlations, and mixed results (e.g., positive revenue impact,…

Applications · Statistics 2026-01-29 Hoiyi Ng , Guido Imbens

Recent work on algorithmic fairness has largely focused on the fairness of discrete decisions, or classifications. While such decisions are often based on risk score models, the fairness of the risk models themselves has received…

Machine Learning · Computer Science 2023-02-24 Eike Petersen , Melanie Ganz , Sune Hannibal Holm , Aasa Feragen

The issue of model risk in default modeling has been known since inception of the Academic literature in the field. However, a rigorous treatment requires a description of all the possible models, and a measure of the distance between a…

Mathematical Finance · Quantitative Finance 2019-06-17 Roberto Fontana , Elisa Luciano , Patrizia Semeraro

We introduce a universal framework for mean-covariance robust risk measurement and portfolio optimization. We model uncertainty in terms of the Gelbrich distance on the mean-covariance space, along with prior structural information about…

Portfolio Management · Quantitative Finance 2025-10-02 Viet Anh Nguyen , Soroosh Shafiee , Damir Filipović , Daniel Kuhn

The standard approach for constructing a Mean-Variance portfolio involves estimating parameters for the model using collected samples. However, since the distribution of future data may not resemble that of the training set, the…

Mathematical Finance · Quantitative Finance 2025-03-12 Duy Khanh Lam

Benchmarks have emerged as the central approach for evaluating Large Language Models (LLMs). The research community often relies on a model's average performance across the test prompts of a benchmark to evaluate the model's performance.…

Computation and Language · Computer Science 2024-06-07 Melissa Ailem , Katerina Marazopoulou , Charlotte Siska , James Bono

Offline reinforcement learning (RL) is suitable for safety-critical domains where online exploration is too costly or dangerous. In such safety-critical settings, decision-making should take into consideration the risk of catastrophic…

Machine Learning · Computer Science 2023-10-31 Marc Rigter , Bruno Lacerda , Nick Hawes

Stochastic portfolio theory aims at finding relative arbitrages, i.e. trading strategies which outperform the market with probability one. Functionally generated portfolios, which are deterministic functions of the market weights, are an…

Mathematical Finance · Quantitative Finance 2021-01-19 Patrick Mijatovic

This paper deals with the scenario approach to robust optimization. This relies on a random sampling of the possibly infinite number of constraints induced by uncertainties in the parameters of an optimization problem. Solving the resulting…

Optimization and Control · Mathematics 2023-03-08 Fabien Lauer

Foundation models are powerful technologies: how they are released publicly directly shapes their societal impact. In this position paper, we focus on open foundation models, defined here as those with broadly available model weights (e.g.…

In this paper, we introduce a probabilistic approach to risk assessment of robot systems by focusing on the impact of uncertainties. While various approaches to identifying systematic hazards (e.g., bugs, design flaws, etc.) can be found in…

Robotics · Computer Science 2024-10-28 Woo-Jeong Baek , Tom P. Huck , Joschka Haas , Jonas Lewandrowski , Tamim Asfour , Torsten Kröger

Large language models often achieve strong benchmark gains without corresponding improvements in broader capability. We hypothesize that this discrepancy arises from differences in training regimes induced by data distribution. To…

Machine Learning · Computer Science 2026-04-10 Hongjian Zou , Yidan Wang , Qi Ding , Yixuan Liao , Xiaoxin Chen

Given the growing influence of language model-based agents on high-stakes societal decisions, from public policy to healthcare, ensuring their beneficial impact requires understanding the far-reaching implications of their suggestions. We…

Artificial Intelligence · Computer Science 2025-06-27 Chenkai Sun , Denghui Zhang , ChengXiang Zhai , Heng Ji

We consider the problem of cost sensitive multiclass classification, where we would like to increase the sensitivity of an important class at the expense of a less important one. We adopt an {\em apportioned margin} framework to address…

Machine Learning · Computer Science 2020-02-05 Lee-Ad Gottlieb , Eran Kaufman , Aryeh Kontorovich