English
Related papers

Related papers: Comparing latent inequality with ordinal data

200 papers

Multivariate categorical data occur in many applications of machine learning. One of the main difficulties with these vectors of categorical variables is sparsity. The number of possible observations grows exponentially with vector length,…

Machine Learning · Statistics 2015-03-10 Yarin Gal , Yutian Chen , Zoubin Ghahramani

A common assumption in causal inference from observational data is that there is no hidden confounding. Yet it is, in general, impossible to verify this assumption from a single dataset. Under the assumption of independent causal mechanisms…

Methodology · Statistics 2023-11-07 Rickard K. A. Karlsson , Jesse H. Krijthe

Organisms and algorithms learn probability distributions from previous observations, either over evolutionary time or on the fly. In the absence of regularities, estimating the underlying distribution from data would require observing each…

Statistical Mechanics · Physics 2024-12-10 William Bialek , Stephanie E. Palmer , David J. Schwab

Real-world time series are influenced by numerous factors and exhibit complex non-stationary characteristics. Non-stationarity can lead to distribution shifts, where the statistical properties of time series change over time, negatively…

Machine Learning · Computer Science 2025-10-13 Zipo Jibao , Yingyi Fu , Xinyang Chen , Guoting Chen

In this paper we focus on providing sufficient conditions for some well-known stochastic orders in reliability but dealing with the discrete versions of them, filling a gap in the literature. In particular, we find conditions based on the…

Statistics Theory · Mathematics 2026-01-28 F. Belzunce , C. Martínez-Riquelme , M. Pereda

Latent space models have been widely adopted in modeling network data. Developing statistical inference for estimated model parameters enables quantifying associated uncertainty and is pivotal for downstream tasks. Despite recent progress…

Statistics Theory · Mathematics 2026-05-12 Yuang Tian , Jiajin Sun , Yinqiu He

Learning a parametric model from a given dataset indeed enables to capture intrinsic dependencies between random variables via a parametric conditional probability distribution and in turn predict the value of a label variable given…

Machine Learning · Statistics 2024-06-14 Elouan Argouarc'h , François Desbouvries , Eric Barat , Eiji Kawasaki

We present a new estimator for causal effects with panel data that builds on insights behind the widely used difference in differences and synthetic control methods. Relative to these methods we find, both theoretically and empirically,…

Methodology · Statistics 2025-09-22 Dmitry Arkhangelsky , Susan Athey , David A. Hirshberg , Guido W. Imbens , Stefan Wager

The presence of unobserved confounders is one of the main challenges in identifying treatment effects. In this paper, we propose a new approach to causal inference using panel data with large large $N$ and $T$. Our approach imputes the…

Econometrics · Economics 2025-03-28 Ben Deaner , Chen-Wei Hsiang , Andrei Zeleneev

Typical causal effects are defined based on the marginal distribution of potential outcomes. However, many real-world applications require causal estimands involving the joint distribution of potential outcomes to enable more nuanced…

Methodology · Statistics 2026-04-17 Peng Wu , Xiaojie Mao

In-context learning (ICL) enables large language models to perform new tasks by conditioning on a sequence of examples. Most prior work reasonably and intuitively assumes that which examples are chosen has a far greater effect on…

Computation and Language · Computer Science 2025-11-14 Warren Li , Yiqian Wang , Zihan Wang , Jingbo Shang

Inference based on the penalized density ratio model is proposed and studied. The model under consideration is specified by assuming that the log--likelihood function of two unknown densities is of some parametric form. The model has been…

Statistics Theory · Mathematics 2008-07-17 Konstantinos Fokianos

Motivated by recently emerging problems in machine learning and statistics, we propose data models which relax the familiar i.i.d. assumption. In essence, we seek to understand what it means for data to come from a set of probability…

Statistics Theory · Mathematics 2025-01-08 Christian Fröhlich , Robert C. Williamson

Two new test statistics are introduced to test the null hypotheses that the sampling distribution has an increasing hazard rate on a specified interval [0,a]. These statistics are empirical L_1-type distances between the isotonic estimates,…

Statistics Theory · Mathematics 2015-03-17 Piet Groeneboom , Geurt Jongbloed

We consider the problem of parameter estimation using weakly supervised datasets, where a training sample consists of the input and a partially specified annotation, which we refer to as the output. The missing information in the annotation…

Machine Learning · Computer Science 2012-06-22 M. Pawan Kumar , Ben Packer , Daphne Koller

Ordinal regression is a classification task where classes have an order and prediction error increases the further the predicted class is from the true class. The standard approach for modeling ordinal data involves fitting parallel…

Machine Learning · Computer Science 2022-02-16 Fred Lu , Francis Ferraro , Edward Raff

Under standard regularity assumptions, we provide simple approximations for specific classes of fiducial and confidence distributions and discuss their connections with objective Bayesian posteriors. For a real parameter the approximations…

Statistics Theory · Mathematics 2017-10-18 Piero Veronese , Eugenio Melilli

Measurement error in observational datasets can lead to systematic bias in inferences based on these datasets. As studies based on observational data are increasingly used to inform decisions with real-world impact, it is critical that we…

Machine Learning · Statistics 2019-01-29 Roy Adams , Yuelong Ji , Xiaobin Wang , Suchi Saria

Databases often contain corrupted, degraded, and noisy data with duplicate entries across and within each database. Such problems arise in citations, medical databases, genetics, human rights databases, and a variety of other applied…

Methodology · Statistics 2015-04-29 Rebecca C. Steorts

Income inequality measures are often used as an indication of economic health. How to obtain reliable confidence intervals for these measures based on sampled data has been studied extensively in recent years. To preserve confidentiality,…

Applications · Statistics 2019-07-22 Dilanka S. Dedduwakumara , Luke A. Prendergast