中文
相关论文

相关论文: Measuring Stochastic Data Complexity with Boltzman…

200 篇论文

Model explainability is crucial for human users to be able to interpret how a proposed classifier assigns labels to data based on its feature values. We study generalized linear models constructed using sets of feature value rules, which…

机器学习 · 统计学 2023-11-06 Sanjeeb Dash , Soumyadip Ghosh , Joao Goncalves , Mark S. Squillante

We show how to achieve the notion of "multicalibration" from H\'ebert-Johnson et al. [2018] not just for means, but also for variances and other higher moments. Informally, it means that we can find regression functions which, given a data…

机器学习 · 计算机科学 2020-08-19 Christopher Jung , Changhwa Lee , Mallesh M. Pai , Aaron Roth , Rakesh Vohra

We investigate the problem of jointly testing a pair of composite hypotheses and, depending on the test result, estimating a random parameter under distributional uncertainties. Specifically, it is assumed that the distribution of the data…

信号处理 · 电气工程与系统科学 2026-04-27 Dominik Reinhard , Michael Fauß , Abdelhak M. Zoubir

Reliable estimation of predictive uncertainty is crucial for machine learning applications, particularly in high-stakes scenarios where hedging against risks is essential. Despite its significance, there is no universal agreement on how to…

机器学习 · 计算机科学 2025-06-17 Kajetan Schweighofer , Lukas Aichberger , Mykyta Ielanskyi , Sepp Hochreiter

Predicting the physico-chemical properties of pure substances and mixtures is a central task in thermodynamics. Established prediction methods range from fully physics-based ab-initio calculations, which are only feasible for very simple…

机器学习 · 计算机科学 2024-12-02 Johannes Zenn , Dominik Gond , Fabian Jirasek , Robert Bamler

In real-world applications, as data availability increases, obtaining labeled data for machine learning (ML) projects remains challenging due to the high costs and intensive efforts required for data annotation. Many ML projects,…

机器学习 · 计算机科学 2024-12-24 Ismail Hakki Karaman , Gulser Koksal , Levent Eriskin , Salih Salihoglu

Calibration ensures that probabilistic forecasts meaningfully capture uncertainty by requiring that predicted probabilities align with empirical frequencies. However, many existing calibration methods are specialized for post-hoc…

机器学习 · 计算机科学 2023-11-01 Charles Marx , Sofian Zalouk , Stefano Ermon

The maximum likelihood estimator in nonlinear panel data models with interactive fixed effects is biased. Several bias correction methods, such as analytical and jackknife approaches, have been proposed to enable valid inference. This paper…

计量经济学 · 经济学 2026-04-30 Haoyuan Xu , Wei Miao , Geert Dhaene , Jad Beyhum

Nonparametric maximum likelihood (NPML) for mixture models is a technique for estimating mixing distributions that has a long and rich history in statistics going back to the 1950s, and is closely related to empirical Bayes methods.…

统计方法学 · 统计学 2018-01-15 Long Feng , Lee H. Dicker

In this paper we present an exploratory research on quantifying the impact that data distribution has on the performance and evaluation of NLP models. We propose an automated framework that measures the data point distribution across 6…

计算与语言 · 计算机科学 2024-04-02 Venelin Kovatchev , Matthew Lease

Methods for split conformal prediction leverage calibration samples to transform any prediction rule into a set-prediction rule that complies with a target coverage probability. Existing methods provide remarkably strong performance…

机器学习 · 统计学 2025-10-15 Santiago Mazuelas

Representing a true label as a one-hot vector is a common practice in training text classification models. However, the one-hot representation may not adequately reflect the relation between the instances and labels, as labels are often not…

计算与语言 · 计算机科学 2020-12-10 Biyang Guo , Songqiao Han , Xiao Han , Hailiang Huang , Ting Lu

A fundamental functional in nonparametric statistics is the Mann-Whitney functional ${\theta} = P (X < Y )$ , which constitutes the basis for the most popular nonparametric procedures. The functional ${\theta}$ measures a location or…

统计方法学 · 统计学 2023-11-30 Jonas Beck , Patrick B. Langthaler , Arne C. Bathke

Quantifying the uncertainty of predictions made by large language models (LLMs) in binary text classification tasks remains a challenge. Calibration, in the context of LLMs, refers to the alignment between the model's predicted…

计算与语言 · 计算机科学 2024-07-02 Patrizio Giovannotti , Alexander Gammerman

Generating confidence calibrated outputs is of utmost importance for the applications of deep neural networks in safety-critical decision-making systems. The output of a neural network is a probability distribution where the scores are…

机器学习 · 计算机科学 2021-09-17 Chihuang Liu , Joseph JaJa

A composite likelihood is an inference function derived by multiplying a set of likelihood components. This approach provides a flexible framework for drawing inference when the likelihood function of a statistical model is computationally…

统计方法学 · 统计学 2024-12-10 Giuseppe Alfonzetti , Ruggero Bellio , Yunxiao Chen , Irini Moustaki

The Normalized Maximum Likelihood (NML) codelength, or stochastic complexity, represents a principled criterion for universal coding. While recent coarea-based formulations provided a calculation method for smooth models, this framework…

机器学习 · 计算机科学 2026-05-26 Trenton Lau , Gary P. T. Choi

Conformal prediction (CP) is a powerful framework for uncertainty quantification, generating prediction sets with coverage guarantees. Split conformal prediction relies on labeled data in the calibration procedure. However, the labeled data…

机器学习 · 计算机科学 2026-03-11 Xuanning Zhou , Zihao Shi , Hao Zeng , Xiaobo Xia , Bingyi Jing , Hongxin Wei

Monitoring machine learning models once they are deployed is challenging. It is even more challenging to decide when to retrain models in real-case scenarios when labeled data is beyond reach, and monitoring performance metrics becomes…

机器学习 · 计算机科学 2022-11-23 Carlos Mougan , Dan Saattrup Nielsen

Empirical best prediction (EBP) is a well-known method for producing reliable proportion estimates when the primary data source provides only small or no sample from finite populations. There are potential challenges in implementing…

统计方法学 · 统计学 2025-01-22 Aditi Sen , Partha Lahiri