English
Related papers

Related papers: Parsimony in Model Selection: Tools for Assessing …

200 papers

The goal of this paper is to compare several widely used Bayesian model selection methods in practical model selection problems, highlight their differences and give recommendations about the preferred approaches. We focus on the variable…

Methodology · Statistics 2017-12-18 Juho Piironen , Aki Vehtari

This paper is dedicated to a robust ordinal method for learning the preferences of a decision maker between subsets. The decision model, derived from Fishburn and LaValle (1996) and whose parameters we learn, is general enough to be…

Artificial Intelligence · Computer Science 2023-08-08 Hugo Gilbert , Mohamed Ouaguenouni , Meltem Ozturk , Olivier Spanjaard

Degrees of freedom is a fundamental concept in statistical modeling, as it provides a quantitative description of the amount of fitting performed by a given procedure. But, despite this fundamental role in statistics, its behavior not…

Statistics Theory · Mathematics 2014-12-01 Ryan J. Tibshirani

Overfitting, which happens when the number of parameters in a model is too large compared to the number of data points available for determining these parameters, is a serious and growing problem in survival analysis. While modern medicine…

Applications · Statistics 2017-09-13 ACC Coolen , JE Barrett , P Paga , CJ Perez-Vicente

In this article, we investigate large sample properties of model selection procedures in a general Bayesian framework when a closed form expression of the marginal likelihood function is not available or a local asymptotic quadratic…

Statistics Theory · Mathematics 2017-01-10 Yun Yang , Debdeep Pati

As data grows in size and complexity, finding frameworks which aid in interpretation and analysis has become critical. This is particularly true when data comes from complex systems where extensive structure is available, but must be drawn…

Machine Learning · Computer Science 2021-05-24 Henry Kvinge , Brett Jefferson , Cliff Joslyn , Emilie Purvine

Choice of appropriate structure and parametric dimension of a model in the light of data has a rich history in statistical research, where the first seminal approaches were developed in 1970s, such as the Akaike's and Schwarz's model…

Methodology · Statistics 2022-06-10 Jukka Corander , Ulpu Remes , Timo Koski

Statistical models for describing the probability distribution over the states of biological systems are commonly used for dimensional reduction. Among these models, pairwise models are very attractive in part because they can be fit using…

Quantitative Methods · Quantitative Biology 2009-11-30 Yasser Roudi , Erik Aurell , John Hertz

In this paper, we examine the problem of building a user profile from a set of documents. This profile will consist of a subset of the most representative terms in the documents that best represent user preferences or interests. Inspired by…

Information Retrieval · Computer Science 2024-01-26 Luis M. de Campos , Juan M. Fernández-Luna , Juan F. Huete

Predictive multiplicity occurs when classification models with statistically indistinguishable performances assign conflicting predictions to individual samples. When used for decision-making in applications of consequence (e.g., lending,…

Machine Learning · Computer Science 2022-10-21 Hsiang Hsu , Flavio du Pin Calmon

Consider the problem of constructing an experimental design, optimal for estimating parameters of a given statistical model with respect to a chosen criterion. To address this problem, the literature usually provides a single solution.…

Computation · Statistics 2024-11-05 Radoslav Harman , Lenka Filová , Samuel Rosa

A common approach in computational science is to use a set of of highly precise but expensive calculations to parameterize a model that allows less precise, but more rapid calculations on larger scale systems. Least-squares fitting on a…

Materials Science · Physics 2015-05-13 Eric Cockayne , Axel van de Walle

The adoption of machine learning in applications where it is crucial to ensure fairness and accountability has led to a large number of model proposals in the literature, largely formulated as optimisation problems with constraints reducing…

Machine Learning · Statistics 2023-05-04 Marco Scutari

The first investigation is made of designs for screening experiments where the response variable is approximated by a generalised linear model. A Bayesian information capacity criterion is defined for the selection of designs that are…

Methodology · Statistics 2016-10-27 David C. Woods , James M. McGree , Susan M. Lewis

Equivariant models leverage prior knowledge on symmetries to improve predictive performance, but misspecified architectural constraints can harm it instead. While work has explored learning or relaxing constraints, selecting among…

Machine Learning · Computer Science 2025-07-16 Putri A. van der Linden , Alexander Timans , Dharmesh Tailor , Erik J. Bekkers

We propose an empirical likelihood test that is able to test the goodness of fit of a class of parametric and semi-parametric multiresponse regression models. The class includes as special cases fully parametric models; semi-parametric…

Statistics Theory · Mathematics 2010-01-12 Song Xi Chen , Ingrid Van Keilegom

The raking-ratio method is a statistical and computational method which adjusts the empirical measure to match the true probability of sets of a finite partition. We study the asymptotic behavior of the raking-ratio empirical process…

Statistics Theory · Mathematics 2019-05-07 Mickael Albertus

Survey sampling is concerned with the estimation of finite population parameters. In practice, survey data suffer from item nonresponse, which is commonly handled through imputation, i.e., replacing missing values with predicted values. As…

Methodology · Statistics 2026-03-06 Ziming An , Mehdi Dagdoug , David Haziza

In the context of regression with a large number of explanatory variables, Cox and Battey (2017) emphasize that if there are alternative reasonable explanations of the data that are statistically indistinguishable, one should aim to specify…

Computation · Statistics 2019-03-15 Henrique Helfer Hoeltgebaum , Heather Battey

The latent class model is a powerful unsupervised clustering algorithm for categorical data. Many statistics exist to test the fit of the latent class model. However, traditional methods to evaluate those fit statistics are not always…

Methodology · Statistics 2018-01-30 Geert H. van Kollenburg , Joris Mulder , Jeroen K. Vermunt