English
Related papers

Related papers: Risk Aware Benchmarking of Large Language Models

200 papers

The classical Markowitz mean-variance model uses variance as a risk measure and calculates frontier portfolios in closed form by using standard optimization techniques. For general mean-risk models such closed form optimal portfolios are…

Mathematical Finance · Quantitative Finance 2026-03-17 Hasanjan Sayit

Both industry and academia have made considerable progress in developing trustworthy and responsible machine learning (ML) systems. While critical concepts like fairness and explainability are often addressed, the safety of systems is…

Machine Learning · Statistics 2022-11-08 Patrick Kaiser , Christoph Kern , David Rügamer

Stochastic dominance serves as a general framework for modeling a broad spectrum of decision preferences under uncertainty, with risk aversion as one notable example, as it naturally captures the intrinsic structure of the underlying…

Machine Learning · Computer Science 2026-01-06 Shicong Cen , Jincheng Mei , Hanjun Dai , Dale Schuurmans , Yuejie Chi , Bo Dai

Evaluation benchmarks are the cornerstone of measuring capabilities of large language models (LLMs), as well as driving progress in said capabilities. Originally designed to make claims about capabilities (or lack thereof) in fully…

Developing large language models is expensive and involves making decisions with small experiments, typically by evaluating on large, multi-task evaluation suites. In this work, we analyze specific properties which make a benchmark more…

Computation and Language · Computer Science 2025-08-19 David Heineman , Valentin Hofmann , Ian Magnusson , Yuling Gu , Noah A. Smith , Hannaneh Hajishirzi , Kyle Lo , Jesse Dodge

We consider the problem of sequential multiple hypothesis testing with nontrivial data collection costs. This problem appears, for example, when conducting biological experiments to identify differentially expressed genes of a disease…

Machine Learning · Computer Science 2023-11-06 Thomas Cook , Harsh Vardhan Dubey , Ji Ah Lee , Guangyu Zhu , Tingting Zhao , Patrick Flaherty

Large Language Models (LLMs) have been widely employed in programming language analysis to enhance human productivity. Yet, their reliability can be compromised by various code distribution shifts, leading to inconsistent outputs. While…

Software Engineering · Computer Science 2024-02-12 Yufei Li , Simin Chen , Yanghong Guo , Wei Yang , Yue Dong , Cong Liu

As the use of machine learning in high impact domains becomes widespread, the importance of evaluating safety has increased. An important aspect of this is evaluating how robust a model is to changes in setting or population, which…

Machine Learning · Computer Science 2021-03-16 Adarsh Subbaswamy , Roy Adams , Suchi Saria

Estimating and assessing the risk of a large portfolio is an important topic in financial econometrics and risk management. The risk is often estimated by a substitution of a good estimator of the volatility matrix. However, the accuracy of…

Applications · Statistics 2013-02-06 Jianqing Fan , Yuan Liao , Xiaofeng Shi

Guard models are a critical component of LLM safety, but their sensitivity to superficial linguistic variations remains a key vulnerability. We show that even meaning-preserving paraphrases can cause large fluctuations in safety scores,…

Computation and Language · Computer Science 2025-11-17 Cristina Pinneri , Christos Louizos

Frontier Large Language Models (LLMs) can be socially discriminatory or sensitive to spurious features of their inputs. Because only well-resourced corporations can train frontier LLMs, we need robust test-time strategies to control such…

Computation and Language · Computer Science 2024-10-08 Leonardo Cotta , Chris J. Maddison

The scope of this manuscript is to review some recent developments in statistics for discretely observed semimartingales which are motivated by applications for financial markets. Our journey through this area stops to take closer looks at…

Statistical Finance · Quantitative Finance 2025-04-23 Markus Bibinger

Motivated by recent work on monotone additive statistics and questions regarding optimal risk sharing for return-based risk measures, we investigate the existence, structure, and applications of Meyer risk measures. Those are monetary risk…

Mathematical Finance · Quantitative Finance 2025-09-30 Christian Laudagé , Felix-Benedikt Liebrich

Large language models are increasingly customized through fine-tuning and other adaptations, creating challenges in enforcing licensing terms and managing downstream impacts. Tracking model origins is crucial both for protecting…

Cryptography and Security · Computer Science 2025-10-31 Ivica Nikolic , Teodora Baluta , Prateek Saxena

This report proposes a novel framework for a rigorous robustness analysis of stochastic biochemical systems. The technique is based on probabilistic model checking. We adapt the general definition of robustness introduced by Kitano to the…

Numerical Analysis · Computer Science 2013-10-18 Lubos Brim , Milan Ceska , Sven Drazan , David Safranek

Understanding variable dependence, particularly eliciting their statistical properties given a set of covariates, provides the mathematical foundation in practical operations management such as risk analysis and decision-making given…

Methodology · Statistics 2023-09-06 Yunyun Wang , Tatsushi Oka , Dan Zhu

The construction of an efficient portfolio with a good level of return and minimal risk depends on selecting the optimal combination of stocks. This paper introduces a novel decision-making framework for stock selection based on fractional…

Statistics Theory · Mathematics 2025-07-04 Poulami Paul , Chanchal Kundu

Searching for new effective risk factors on stock returns is an important research topic in asset pricing. Factor modeling is an active research topic in statistics and econometrics, with many new advances. However, these new methods have…

Risk Management · Quantitative Finance 2024-09-27 Xialu Liu , John Guerard , Rong Chen , Ruey Tsay

The growing adoption of large language models (LLMs) in finance exposes high-stakes decision-making to subtle, underexamined positional biases. The complexity and opacity of modern model architectures compound this risk. We present the…

Computational Finance · Quantitative Finance 2025-10-08 Fabrizio Dimino , Krati Saxena , Bhaskarjit Sarmah , Stefano Pasquali

Fairness-aware statistical learning is essential for mitigating discrimination against protected attributes such as gender, race, and ethnicity in data-driven decision-making. This is particularly critical in high-stakes applications like…

Methodology · Statistics 2025-04-15 Fei Huang , Junhao Shen , Yanrong Yang , Ran Zhao