English
Related papers

Related papers: Pseudo-$R^2$ statistics under complex sampling

200 papers

This paper is concerned with sample size determination methodology for prediction models. We propose combining the individual calculations via a learning-type curve. We suggest two distinct ways of doing so, a deterministic skeleton of a…

Methodology · Statistics 2024-05-24 Alimu Dayimu , Nikola Simidjievski , Nikolaos Demiris , Jean Abraham

The use of big data in official statistics and the applied sciences is accelerating, but statistics computed using only big data often suffer from substantial selection bias. This leads to inaccurate estimation and invalid statistical…

Methodology · Statistics 2023-08-11 Ryan Covey , Lucca Buonamano

Modern applications require methods that are computationally feasible on large datasets but also preserve statistical efficiency. Frequently, these two concerns are seen as contradictory: approximation methods that enable computation are…

Methodology · Statistics 2021-06-11 Darren Homrighausen , Daniel J. McDonald

Generalized linear models (GLMs) -- such as logistic regression, Poisson regression, and robust regression -- provide interpretable models for diverse data types. Probabilistic approaches, particularly Bayesian ones, allow coherent…

Computation · Statistics 2018-12-19 Jonathan H. Huggins , Ryan P. Adams , Tamara Broderick

We propose a generalisation of the logistic regression model, that aims to account for non-linear main effects and complex interactions, while keeping the model inherently explainable. This is obtained by starting with log-odds that are…

Methodology · Statistics 2024-10-14 Ingrid Hobæk Haff , Simon Boge Brant , Haakon Bakka

When cross-validating standard or extended Cox models, the commonly used criterion is the cross-validated partial loglikelihood using a naive or a van Houwelingen scheme -to make efficient use of the death times of the left out data in…

Methodology · Statistics 2018-10-09 Frédéric Bertrand , Philippe Bastien , Myriam Maumy-Bertrand

Linear regression is a fundamental building block of statistical data analysis. It amounts to estimating the parameters of a linear model that maps input features to corresponding outputs. In the classical setting where the precision of…

Computer Science and Game Theory · Computer Science 2019-12-16 Nicolas Gast , Stratis Ioannidis , Patrick Loiseau , Benjamin Roussillon

By means of two simple convexity arguments we are able to develop a general method for proving consistency and asymptotic normality of estimators that are defined by minimisation of convex criterion functions. This method is then applied to…

Statistics Theory · Mathematics 2011-07-20 Nils Lid Hjort , David Pollard

In a Cox model, the partial likelihood, as the product of a series of conditional probabilities, is used to estimate the regression coefficients. In practice, those conditional probabilities are approximated by risk score ratios based on a…

Methodology · Statistics 2025-02-27 Youngjin Cho , Yili Hong , Pang Du

Deep sequence models are receiving significant interest in current machine learning research. By representing probability distributions that are fit to data using maximum likelihood estimation, such models can model data on general…

Systems and Control · Electrical Eng. & Systems 2024-09-09 Kristian Løvland , Bjarne Grimstad , Lars Struen Imsland

A general structural equation model is fitted on a panel data set that consists of $I$ correlated samples. The correlated samples could be data from correlated populations or correlated observations from occasions of panel data. We consider…

Statistics Theory · Mathematics 2007-06-13 Savas Papadopoulos , Yasuo Amemiya

To generalize inferences from a randomized trial to the target population of all trial-eligible individuals, investigators can use nested trial designs, where the randomized individuals are nested within a cohort of trial-eligible…

This article investigates uncertainty quantification of the generalized linear lasso~(GLL), a popular variable selection method in high-dimensional regression settings. In many fields of study, researchers use data-driven methods to select…

Statistics Theory · Mathematics 2023-07-11 Quentin Duchemin , Yohann de Castro

Simulation-based inference plays a major role in modern statistics, and often employs either reallocating (as in a randomization test) or resampling (as in bootstrapping). Reallocating mimics random allocation to treatment groups, while…

Statistics Theory · Mathematics 2017-08-08 Kari Lock Morgan

The cost of both generalized least squares (GLS) and Gibbs sampling in a crossed random effects model can easily grow faster than $N^{3/2}$ for $N$ observations. Ghosh et al. (2020) develop a backfitting algorithm that reduces the cost to…

Methodology · Statistics 2021-12-30 Swarnadip Ghosh , Trevor Hastie , Art B. Owen

We propose a generalized functional linear regression model for a regression situation where the response variable is a scalar and the predictor is a random function. A linear predictor is obtained by forming the scalar product of the…

Statistics Theory · Mathematics 2007-06-13 Hans-Georg Muller , Ulrich Stadtmuller

The analysis of non-real-valued data, such as binary time series, has attracted great interest in recent years. This manuscript proposes a post-selection estimator for estimating the coefficient matrices of a high-dimensional generalized…

Methodology · Statistics 2025-12-03 Dehao Dai , Yunyi Zhang

Downsampling or under-sampling is a technique that is utilized in the context of large and highly imbalanced classification models. We study optimal downsampling for imbalanced classification using generalized linear models (GLMs). We…

Machine Learning · Statistics 2025-05-20 Yan Chen , Jose Blanchet , Krzysztof Dembczynski , Laura Fee Nern , Aaron Flores

We consider the problem of statistical inference for ranking data, specifically rank aggregation, under the assumption that samples are incomplete in the sense of not comprising all choice alternatives. In contrast to most existing methods,…

Machine Learning · Statistics 2017-12-05 Mohsen Ahmadi Fahandar , Eyke Hüllermeier , Inés Couso

In this paper it is shown that under certain conditions there is a relationship between the parameter estimation of the Fellegi--Sunter probabilistic linkage model and dual system estimation. This relationship can be used as the basis of an…

Methodology · Statistics 2019-03-27 Viktor Račinskij , Paul A. Smith , Peter G. M. van der Heijden