English
Related papers

Related papers: Addressing the Impact of Data Truncation and Param…

200 papers

Survival analysis is a fundamental tool for modeling time-to-event data in healthcare, engineering, and finance, where censored observations pose significant challenges. While traditional methods like the Beran estimator offer nonparametric…

Machine Learning · Computer Science 2025-06-13 Andrei V. Konstantinov , Vlada A. Efremenko , Lev V. Utkin

Recent empirical and theoretical analyses of several commonly used prediction procedures reveal a peculiar risk behavior in high dimensions, referred to as double/multiple descent, in which the asymptotic risk is a non-monotonic function of…

Statistics Theory · Mathematics 2022-05-26 Pratik Patil , Arun Kumar Kuchibhotla , Yuting Wei , Alessandro Rinaldo

Missing values arise in most real-world data sets due to the aggregation of multiple sources and intrinsically missing information (sensor failure, unanswered questions in surveys...). In fact, the very nature of missing values usually…

Machine Learning · Statistics 2022-02-04 Alexis Ayme , Claire Boyer , Aymeric Dieuleveut , Erwan Scornet

Uncertainty is a key feature of any machine learning model and is particularly important in neural networks, which tend to be overconfident. This overconfidence is worrying under distribution shifts, where the model performance silently…

Machine Learning · Computer Science 2024-03-18 Arthur Thuy , Dries F. Benoit

We propose a truncation model for abundance distribution in the species richness estimation. This model is inherently semiparametric and incorporates an unknown truncation threshold between rare and abundant counts observations. Using the…

Methodology · Statistics 2017-05-23 François Koladjo , Mesrob I. Ohannessian , Élisabeth Gassiat

Analysis of longitudinal randomised controlled trials is frequently complicated because patients deviate from the protocol. Where such deviations are relevant for the estimand, we are typically required to make an untestable assumption…

Methodology · Statistics 2018-05-16 Suzie Cro , James R Carpenter , Michael G Kenward

The majority of existing probabilistic model checking case studies are based on well understood theoretical models and distributions. However, real-life probabilistic systems usually involve distribution parameters whose values are obtained…

Software Engineering · Computer Science 2013-08-29 Guoxin Su , David S. Rosenblum

We study high-confidence off-policy evaluation in the context of infinite-horizon Markov decision processes, where the objective is to establish a confidence interval (CI) for the target policy value using only offline data pre-collected…

Machine Learning · Statistics 2023-10-03 Wenzhuo Zhou , Yuhan Li , Ruoqing Zhu , Annie Qu

This paper introduces a high frequency trade execution model to evaluate the economic impact of supervised machine learners. Extending the concept of a confusion matrix, we present a 'trade information matrix' to attribute the expected…

Trading and Market Microstructure · Quantitative Finance 2017-12-06 Matthew F Dixon

When data are missing due to at most one cause from some time to next time, we can make sampling distribution inferences about the parameter of the data by modeling the missing-data mechanism correctly. Proverbially, in case its mechanism…

Methodology · Statistics 2014-07-21 Kosuke Morikawa , Yutaka Kano

Memorization in over-parameterized neural networks could severely hurt generalization in the presence of mislabeled examples. However, mislabeled examples are hard to avoid in extremely large datasets collected with weak supervision. We…

Machine Learning · Computer Science 2020-04-10 Jiaming Song , Lunjia Hu , Michael Auli , Yann Dauphin , Tengyu Ma

Conformal unlearning aims to ensure that a trained conformal predictor miscovers data points with specific shared characteristics, such as those from a particular label class, associated with a specific user, or belonging to a defined…

Machine Learning · Computer Science 2026-02-13 Yahya Alkhatib , Muhammad Ahmar Jamal , Wee Peng Tay

We posit that data can only be safe to use up to a certain threshold of the data distribution shift, after which control must be relinquished by the autonomous system and operation halted or handed to a human operator. With the use of a…

Machine Learning · Computer Science 2024-07-01 Daniel Sikar , Artur Garcez

The stochastic logistic model with regime switching is an important model in the ecosystem. While analytic solution to this model is positive, current numerical methods are unable to preserve such boundaries in the approximation. So,…

Numerical Analysis · Mathematics 2021-06-08 Xiaoyue Li , Hongfu Yang

Computing partition functions, the normalizing constants of probability distributions, is often hard. Variants of importance sampling give unbiased estimates of a normalizer Z, however, unbiased estimates of the reciprocal 1/Z are harder to…

Machine Learning · Statistics 2017-03-14 Colin Wei , Iain Murray

Link prediction is a fundamental problem in network science, aiming to infer potential or missing links based on observed network structures. With the increasing adoption of parameterized models, the rigor of evaluation protocols has become…

Other Statistics · Statistics 2026-04-09 Xinshan Jiao , Yuxin Luo , Yilin Bi , Tao Zhou

Estimating the effects of continuous-valued interventions from observational data is a critically important task for climate science, healthcare, and economics. Recent work focuses on designing neural network architectures and…

Machine Learning · Computer Science 2022-10-13 Andrew Jesson , Alyson Douglas , Peter Manshausen , Maëlys Solal , Nicolai Meinshausen , Philip Stier , Yarin Gal , Uri Shalit

Learning models or control policies from data has become a powerful tool to improve the performance of uncertain systems. While a strong focus has been placed on increasing the amount and quality of data to improve performance, data can…

Systems and Control · Electrical Eng. & Systems 2024-10-02 Ralf Römer , Lukas Brunke , Siqi Zhou , Angela P. Schoellig

When a source-trained model $Q$ is replaced by a model $\tilde{Q}$ trained on shifted data, its performance on the source domain can change unpredictably. To address this, we study the two-model risk change, $\Delta R := R_P(\tilde{Q}) -…

Machine Learning · Computer Science 2026-02-12 Hosein Anjidani , S. Yahya S. R. Tehrani , Mohammad Mahdi Mojahedian , Mohammad Hossein Yassaee

A conventional linear model for functional data involves expressing a response variable $Y$ in terms of the explanatory function $X(t)$, via the model: $Y=a+\int_I b(t)X(t)dt+\hbox{error}$, where $a$ is a scalar, $b$ is an unknown function…

Methodology · Statistics 2014-07-01 Peter Hall , Giles Hooker