English
Related papers

Related papers: On Modeling Profiles instead of Values

200 papers

Distributions following a power-law are an ubiquitous phenomenon. Methods for determining the exponent of a power-law tail by graphical means are often used in practice but are intrinsically unreliable. Maximum likelihood estimators for the…

Other Condensed Matter · Physics 2007-08-11 Heiko Bauke

Estimation frameworks for statistical inference are preferred to hypothesis testing when quantifying uncertainty and precise estimation are more valuable than binary decisions about statistical significance. Study design for…

Methodology · Statistics 2025-10-29 Luke Hagar , Nathaniel T. Stevens

We present an algorithmic approach to estimate the value distributions of random variables of probabilistic loops whose statistical moments are (partially) known. Based on these moments, we apply two statistical methods, Maximum Entropy and…

Symbolic data analysis (SDA) is an emerging area of statistics concerned with understanding and modelling data that takes distributional form (i.e. symbols), such as random lists, intervals and histograms. It was developed under the premise…

Computation · Statistics 2020-04-09 Boris Beranger , Huan Lin , Scott A. Sisson

Maximum-likelihood exponent maps have been studied as a technique to increase the understanding and improve the fit of power-law exponents to experimental and numerical simulation data, especially when they exhibit both upper and lower…

Statistical Mechanics · Physics 2012-07-02 Jordi Baró , Eduard Vives

Recent likelihood theory produces $p$-values that have remarkable accuracy and wide applicability. The calculations use familiar tools such as maximum likelihood values (MLEs), observed information and parameter rescaling. The usual…

Methodology · Statistics 2008-02-08 M. Bédard , D. A. S. Fraser , A. Wong

In numerous instances, the generalized exponential distribution can be used as an alternative to the most widely used non-regular family of distributions: Weibull, gamma, lognormal with three-parameters when analyzing lifetime or any skewed…

Methodology · Statistics 2026-03-03 Kiran Prajapat , Sharmishtha Mitra , Debasis Kundu

Suppose that we are given a time series where consecutive samples are believed to come from a probabilistic source, that the source changes from time to time and that the total number of sources is fixed. Our objective is to estimate the…

Information Theory · Computer Science 2018-04-24 Mark Kozdoba , Shie Mannor

Modelling human variation in rating tasks is crucial for personalization, pluralistic model alignment, and computational social science. We propose representing individuals using natural language value profiles -- descriptions of underlying…

It is impossible today to pretend that the practice of machine learning is always compatible with the idea that training and testing data follow the same distribution. Several authors have recently used ensemble techniques to show how…

Machine Learning · Computer Science 2025-03-03 Jianyu Zhang , Léon Bottou

We consider the problem of estimating the probability of an observed string drawn i.i.d. from an unknown distribution. The key feature of our study is that the length of the observed string is assumed to be of the same order as the size of…

Information Theory · Computer Science 2007-07-13 Aaron B. Wagner , Pramod Viswanath , Sanjeev R. Kulkarni

In this article we discuss estimation of the common variance of several normal populations with tree order restricted means. We discuss the asymptotic properties of the maximum likelihood estimator of the variance as the number of…

Statistics Theory · Mathematics 2014-07-24 Antar Bandyopadhyay , Sanjay Chaudhuri

Paired comparison data considered in this paper originate from the comparison of a large number N of individuals in couples. The dataset is a collection of results of contests between two individuals when each of them has faced n opponents,…

Statistics Theory · Mathematics 2020-02-14 Roland Diel , Sylvain Le Corff , Matthieu Lerasle

Counts of attribute-value combinations are central to the profiling of a dataset, particularly in determining fitness for use and in eliminating bias and unfairness. While counts of individual attribute values may be stored in some dataset…

Databases · Computer Science 2020-11-10 Yuval Moskovitch , H. V. Jagadish

In a multifidelity setting, data are available under the same conditions from two (or more) sources, e.g. computer codes, one being lower-fidelity but computationally cheaper, and the other higher-fidelity and more expensive. This work…

Methodology · Statistics 2024-09-25 Minji Kim , Kevin O'Connor , Vladas Pipiras , Themistoklis Sapsis

The blessing of ubiquitous data also comes with a curse: the communication, storage, and labeling of massive, mostly redundant datasets. We seek to solve this problem at its core, collecting only valuable data and throwing out the rest via…

Machine Learning · Computer Science 2023-12-18 Mariel Werner , Anastasios Angelopoulos , Stephen Bates , Michael I. Jordan

F\'elix-Medina and Thompson (2004) proposed a variant of link-tracing sampling to estimate the size of a hidden population such as drug users, sexual workers or homeless people. In their variant a sampling frame of sites where the members…

Methodology · Statistics 2015-06-23 Martin H. Félix Medina

In multivariate or spatial extremes, inference for max-stable processes observed at a large collection of locations is among the most challenging problems in computational statistics, and current approaches typically rely on less expensive…

Computation · Statistics 2015-08-20 Stefano Castruccio , Raphaël Huser , Marc Genton

In this short note, we derive a new bias adjusted maximum likelihood estimate for the shape parameter of the Weibull distribution with complete data and type I censored data. The proposed estimate of the shape parameter is significantly…

Methodology · Statistics 2023-02-13 Enes Makalic , Daniel F. Schmidt

When analyzing incomplete data, is it better to use multiple imputation (MI) or full information maximum likelihood (ML)? In large samples ML is clearly better, but in small samples ML's usefulness has been limited because ML commonly uses…

Methodology · Statistics 2017-03-24 Paul T. von Hippel