English
Related papers

Related papers: Cautionary note on "Semiparametric modeling of gro…

200 papers

Recent language models generate false but plausible-sounding text with surprising frequency. Such "hallucinations" are an obstacle to the usability of language-based AI systems and can harm people who rely upon their outputs. This work…

Computation and Language · Computer Science 2024-03-21 Adam Tauman Kalai , Santosh S. Vempala

The analysis of randomized trials with time-to-event endpoints is nearly always plagued by the problem of censoring. As the censoring mechanism is usually unknown, analyses typically employ the assumption of non-informative censoring. While…

Methodology · Statistics 2020-07-17 Kelly Van Lancker , Oliver Dukes , Stijn Vansteelandt

Causal inference necessarily relies upon untestable assumptions; hence, it is crucial to assess the robustness of obtained results to violations of identification assumptions. However, such sensitivity analysis is only occasionally…

Methodology · Statistics 2025-05-19 Tobias Freidling , Qingyuan Zhao

In medical settings, treatment assignment may be determined by a clinically important covariate that predicts patients' risk of event. There is a class of methods from the social science literature known as regression discontinuity (RD)…

Methodology · Statistics 2019-08-13 Youngjoo Cho , Chen Hu , Debashis Ghosh

To get Bayesian neural networks to perform comparably to standard neural networks it is usually necessary to artificially reduce uncertainty using a "tempered" or "cold" posterior. This is extremely concerning: if the prior is accurate,…

Machine Learning · Statistics 2021-04-28 Laurence Aitchison

The limit distribution of the nonparametric maximum likelihood estimator for interval censored data with more than one observation time per unobservable observation, is still unknown in general. For the so-called separated case, where one…

Statistics Theory · Mathematics 2026-02-12 Piet Groeneboom

In the recent Bayesian nonparametric literature, many examples have been reported in which Bayesian estimators and posterior distributions do not achieve the optimal convergence rate, indicating that the Bernstein-von Mises theorem does not…

Statistics Theory · Mathematics 2007-06-13 Yongdai Kim , Jaeyong Lee

Imputation is a popular approach to handling censored, missing, and error-prone covariates -- all coarsened data types for which the true values are unknown. However, there are nuances to imputing these different data types based on the…

Methodology · Statistics 2025-04-29 Sarah C. Lotspeich , Ethan M. Alt

We discuss the role of misspecification and censoring on Bayesian model selection in the contexts of right-censored survival and concave log-likelihood regression. Misspecification includes wrongly assuming the censoring mechanism to be…

Methodology · Statistics 2021-11-15 David Rossell , Francisco Javier Rubio

Recent works in artificial intelligence fairness attempt to mitigate discrimination by proposing constrained optimization programs that achieve parity for some fairness statistic. Most assume availability of the class label, which is…

Machine Learning · Computer Science 2022-04-01 Wenbin Zhang , Jeremy C. Weiss

The i.i.d. censoring model for survival analysis assumes two independent sequences of i.i.d. positive random variables, $(T_i^*)_{1\le i\le n}$ and $(U_i)_{1\le i\le n}$. The data consists of observations on the random sequence…

Statistics Theory · Mathematics 2020-02-27 Ross A. Maller , Sidney I. Resnick

Estimating the prevalence of a category in a population using imperfect measurement devices (diagnostic tests, classifiers, or large language models) is fundamental to science, public health, and online trust and safety. Standard approaches…

Artificial Intelligence · Computer Science 2026-04-24 Fridolin Linder , Thomas Leeper , Daniel Haimovich , Niek Tax , Lorenzo Perini , Milan Vojnovic

Unsupervised learning is often used to uncover clusters in data. However, different kinds of noise may impede the discovery of useful patterns from real-world time-series data. In this work, we focus on mitigating the interference of…

Machine Learning · Statistics 2021-12-07 Irene Y. Chen , Rahul G. Krishnan , David Sontag

This article proposes inference procedures for distribution regression models in duration analysis using randomly right-censored data. This generalizes classical duration models by allowing situations where explanatory variables' marginal…

Econometrics · Economics 2021-11-29 Miguel A. Delgado , Andrés García-Suaza , Pedro H. C. Sant'Anna

As machine learning methods are deployed in real-world settings such as healthcare, legal systems, and social science, it is crucial to recognize how they shape social biases and stereotypes in these sensitive decision-making processes.…

Computation and Language · Computer Science 2021-06-25 Paul Pu Liang , Chiyu Wu , Louis-Philippe Morency , Ruslan Salakhutdinov

Social biases such as gender or racial biases have been reported in language models (LMs), including Masked Language Models (MLMs). Given that MLMs are continuously trained with increasing amounts of additional data collected over time, an…

Computation and Language · Computer Science 2024-06-21 Yi Zhou , Danushka Bollegala , Jose Camacho-Collados

In survival studies, classical inferences for left-truncated data require quasi-independence, a property that the joint density of truncation time and failure time is factorizable into their marginal densities in the observable region. The…

Methodology · Statistics 2019-04-16 Young-Geun Choi , Wei-Yann Tsai , Myunghee Cho Paik

Tailoring treatments to individual needs is a central goal in fields such as medicine. A key step toward this goal is estimating Heterogeneous Treatment Effects (HTE) - the way treatments impact different subgroups. While crucial, HTE…

Machine Learning · Statistics 2025-07-30 Tomer Meir , Uri Shalit , Malka Gorfine

Large Language Models (LLMs) are increasingly used for synthetic tabular data generation through in-context learning (ICL), offering a practical solution for data augmentation in data scarce scenarios. While prior work has shown the…

The Weibull distribution is one of the most used tools in reliability analysis. In this paper, assuming a Bayesian approach, we propose necessary and sufficient conditions to verify when improper priors lead to proper posteriors for the…

Statistics Theory · Mathematics 2020-05-19 Eduardo Ramos , Pedro L. Ramos