English
Related papers

Related papers: Codage arithmetique pour la description d'une dist…

200 papers

We transpose an optimal control technique to the image segmentation problem. The idea is to consider image segmentation as a parameter estimation problem. The parameter to estimate is the color of the pixels of the image. We use the…

Numerical Analysis · Mathematics 2011-05-24 Hend Ben Ameur , Guy Chavent , Francois Clément , Pierre Weis

The first investigation is made of designs for screening experiments where the response variable is approximated by a generalised linear model. A Bayesian information capacity criterion is defined for the selection of designs that are…

Methodology · Statistics 2016-10-27 David C. Woods , James M. McGree , Susan M. Lewis

A central issue of many statistical learning problems is to select an appropriate model from a set of candidate models. Large models tend to inflate the variance (or overfitting), while small models tend to cause biases (or underfitting)…

Statistics Theory · Mathematics 2020-12-25 Jie Ding , Enmao Diao , Jiawei Zhou , Vahid Tarokh

We propose information criteria that measure the prediction risk of a predictive density based on the Bayesian marginal likelihood from a frequentist point of view. We derive criteria for selecting variables in linear regression models,…

Methodology · Statistics 2017-10-20 Yuki Kawakubo , Tatsuya Kubokawa , Muni S. Srivastava

Random network models, constrained to reproduce specific statistical features, are often used to represent and analyze network data and their mathematical descriptions. Chief among them, the configuration model constrains random networks by…

Social and Information Networks · Computer Science 2025-01-28 Laurent Hébert-Dufresne , Jean-Gabriel Young , Alexander Daniels , Alec Kirkley , Antoine Allard

The non-identifiability of the competing risks model requires researchers to work with restrictions on the model to obtain informative results. We present a new identifiability solution based on an exclusion restriction. Many areas of…

Methodology · Statistics 2023-09-06 Munir Hiabu , Simon M. S. LU , Ralf A. Wilke

This article considers the problem of modeling a class of nonstationary count time series using multiple change-points generalized integer-valued autoregressive (MCP-GINAR) processes. The minimum description length principle (MDL) is…

Applications · Statistics 2023-07-04 Danshu Sheng , Dehui Wang

Shi and Tsai (JRSSB, 2002) proposed an interesting residual information criterion (RIC) for model selection in regression. Their RIC was motivated by the principle of minimizing the Kullback-Leibler discrepancy between the residual…

Methodology · Statistics 2007-11-14 Chenlei Leng

Conventional likelihood-based information criteria for model selection rely on the distribution assumption of data. However, for complex data that are increasingly available in many scientific fields, the specification of their underlying…

Methodology · Statistics 2020-06-25 Chixiang Chen , Ming Wang , Rongling Wu , Runze Li

Data compression often subtracts prediction and encodes the difference (residue) e.g. assuming Laplace distribution, for example for images, videos, audio, or numerical data. Its performance is strongly dependent on the proper choice of…

Image and Video Processing · Electrical Eng. & Systems 2019-10-15 Jarek Duda

We consider the classical problem of discrete distribution estimation using i.i.d. samples in a novel scenario where additional side information is available on the distribution. In large alphabet datasets such as text corpora, such side…

Information Theory · Computer Science 2026-01-19 Haricharan Balasundaram , Andrew Thangaraj

Many important modeling tasks in linear regression, including variable selection (in which slopes of some predictors are set equal to zero) and simplified models based on sums or differences of predictors (in which slopes of those…

Methodology · Statistics 2020-09-22 Sen Tian , Clifford M. Hurvich , Jeffrey S. Simonoff

Linear discriminant analysis is a widely used method for classification. However, the high dimensionality of predictors combined with small sample sizes often results in large classification errors. To address this challenge, it is crucial…

Machine Learning · Statistics 2025-01-09 Hongzhe Zhang , Arnab Auddy , Hongzhe Lee

Divergence From Randomness (DFR) ranking models assume that informative terms are distributed in a corpus differently than non-informative terms. Different statistical models (e.g. Poisson, geometric) are used to model the distribution of…

Information Retrieval · Computer Science 2016-09-06 Casper Petersen , Jakob Grue Simonsen , Kalervo Jarvelin , Christina Lioma

The main contribution of this paper is to design an Information Retrieval (IR) technique based on Algorithmic Information Theory (using the Normalized Compression Distance- NCD), statistical techniques (outliers), and novel organization of…

Information Retrieval · Computer Science 2016-11-17 Rafael Martinez , Manuel Cebrian , Francisco de Borja Rodriguez , David Camacho

For parameterized mixed-binary optimization problems, we construct local decision rules that prescribe near-optimal courses of action across a set of parameter values. The decision rules stem from solving risk-adaptive training problems…

Optimization and Control · Mathematics 2024-04-24 Johannes O. Royset , Miguel A. Lejeune

Unsupervised discretization is a crucial step in many knowledge discovery tasks. The state-of-the-art method for one-dimensional data infers locally adaptive histograms using the minimum description length (MDL) principle, but the…

Machine Learning · Computer Science 2022-12-12 Lincen Yang , Mitra Baratchi , Matthijs van Leeuwen

Data quality is critical for multimedia tasks, while various types of systematic flaws are found in image benchmark datasets, as discussed in recent work. In particular, the existence of the semantic gap problem leads to a many-to-many…

Computer Vision and Pattern Recognition · Computer Science 2023-04-19 Fausto Giunchiglia , Xiaolei Diao , Mayukh Bagchi

A popular technique for selecting and tuning machine learning estimators is cross-validation. Cross-validation evaluates overall model fit, usually in terms of predictive accuracy. In causal inference, the optimal choice of estimator…

Methodology · Statistics 2021-07-07 Dominik Rothenhäusler

Log-linear models are a well-established method for describing statistical dependencies among a set of n random variables. The observed frequencies of the n-tuples are explained by a joint probability such that its logarithm is a sum of…

Statistics Theory · Mathematics 2007-06-13 Daniel Herrmann , Dominik Janzing
‹ Prev 1 8 9 10 Next ›