English
Related papers

Related papers: Plotting the Differences Between Data and Expectat…

200 papers

Statistical systems are conceived from the standpoint of statistical mechanics, as made of a (generally large) number of identical units and exhibiting a (generally large) number of different configurations (microstates), among which only…

General Physics · Physics 2017-06-21 R. Caimmi

Logistic regression models are a popular and effective method to predict the probability of categorical response data. However inference for these models can become computationally prohibitive for large datasets. Here we adapt ideas from…

Methodology · Statistics 2020-08-25 Tom Whitaker , Boris Beranger , Scott A. Sisson

Nonlinear expectation, including sublinear expectation as its special case, is a new and original framework of probability theory and has potential applications in some scientific fields, especially in finance risk measure and management.…

Statistics Theory · Mathematics 2013-04-15 Lu Lin , Yufeng Shi , Xin Wang , Shuzhen Yang

Visual statistical inference is a way to determine significance of patterns found while exploring data. It is dependent on the evaluation of a lineup, of a data plot among a sample of null plots, by human observers. Each individual is…

Applications · Statistics 2014-08-12 Mahbubul Majumder , Heike Hofmann , Dianne Cook

Imputing missing values is an important preprocessing step in data analysis, but the literature offers little guidance on how to choose between different imputation models. This letter suggests adopting the imputation model that generates a…

Methodology · Statistics 2021-07-13 Moritz Marbach

Given a large dataset and an estimation task, it is common to pre-process the data by reducing them to a set of sufficient statistics. This step is often regarded as straightforward and advantageous (in that it simplifies statistical…

Computation · Statistics 2015-07-31 Andrea Montanari

The histogram method is a powerful non-parametric approach for estimating the probability density function of a continuous variable. But the construction of a histogram, compared to the parametric approaches, demands a large number of…

Machine Learning · Statistics 2015-12-29 Hideaki Kim , Hiroshi Sawada

When the data do not conform to the hypothesis of a known sampling-variance, the fitting of a constant to a set of measured values is a long debated problem. Given the data, fitting would require to find what measurand value is the most…

Data Analysis, Statistics and Probability · Physics 2020-07-21 Giovanni Mana , Enrico Massa , Maria Predescu

Density plots effectively summarize large numbers of points, which would otherwise lead to severe overplotting in, for example, a scatter plot. However, when applied to line-based datasets, such as trajectories or time series, density plots…

Graphics · Computer Science 2025-12-19 Yumeng Xue , Bin Chen , Patrick Paetzold , Yunhai Wang , Christophe Hurter , Oliver Deussen

Consistent experiment data are crucial to adjust parameters of physics models and to determine best estimates of observables. However, often experiment data are not consistent due to unrecognized systematic errors. Standard methods of…

Nuclear Theory · Physics 2018-03-05 Georg Schnabel

The recent explosion in the amount and dimensionality of data has exacerbated the need of trading off computational and statistical efficiency carefully, so that inference is both tractable and meaningful. We propose a framework that…

Computation · Statistics 2015-06-29 Daniel L. Sussman , Alexander Volfovsky , Edoardo M. Airoldi

Expected shortfall is defined as the average over the tail below (or above) a certain quantile of a probability distribution. Expected shortfall regression provides powerful tools for learning the relationship between a response variable…

Methodology · Statistics 2025-01-03 Shushu Zhang , Xuming He , Kean Ming Tan , Wen-Xin Zhou

Redundancy of experimental data is the basic statistic from which the complexity of a natural phenomenon and the proper number of experiments needed for its exploration can be estimated. The redundancy is expressed by the entropy of…

Data Analysis, Statistics and Probability · Physics 2007-10-10 I. Grabec

The bagplot, also known as the "bag-and-bolster plot", is a notable extension of the boxplot from univariate to bivariate data. Although widely used, its practical application is hindered by two key limitations: the fixed inflation factor…

Methodology · Statistics 2025-12-09 Shenghao Qin , Bowen Gang , Tiejun Tong , Hengjian Cui

Making binary decisions is a common data analytical task in scientific research and industrial applications. In data sciences, there are two related but distinct strategies: hypothesis testing and binary classification. In practice, how to…

Applications · Statistics 2021-12-01 Jingyi Jessica Li , Xin Tong

In many analyses in high energy physics, attempts are made to remove the effects of detector smearing in data by techniques referred to as "unfolding" histograms, thus obtaining estimates of the true values of histogram bin contents. Such…

Data Analysis, Statistics and Probability · Physics 2016-07-26 Robert D. Cousins , Samuel J. May , Yipeng Sun

We consider basic conceptual questions concerning the relationship between statistical estimation and causal inference. Firstly, we show how to translate causal inference problems into an abstract statistical formalism without requiring any…

Statistics Theory · Mathematics 2020-07-22 Oliver J. Maclaren , Ruanui Nicholson

A research frontier has emerged in scientific computation, wherein numerical error is regarded as a source of epistemic uncertainty that can be modelled. This raises several statistical challenges, including the design of statistical…

Machine Learning · Statistics 2017-10-19 François-Xavier Briol , Chris. J. Oates , Mark Girolami , Michael A. Osborne , Dino Sejdinovic

Estimation of the $\phi$-divergence between two unknown probability distributions using empirical data is a fundamental problem in information theory and statistical learning. We consider a multi-variate generalization of the data dependent…

Probability · Mathematics 2018-01-04 Fengqiao Luo , Sanjay Mehrotra

We propose a new measure of deviations from expected utility theory. For any positive number~$e$, we give a characterization of the datasets with a rationalization that is within~$e$ (in beliefs, utility, or perceived prices) of expected…

General Economics · Economics 2021-02-15 Federico Echenique , Kota Saito , Taisuke Imai