English
Related papers

Related papers: Matrix dissimilarities based on differences in mom…

200 papers

The performance of standard learning procedures has been observed to differ widely across groups. Recent studies usually attribute this loss discrepancy to an information deficiency for one group (e.g., one group has less data). In this…

Machine Learning · Computer Science 2020-11-09 Fereshte Khani , Percy Liang

With the widespread deployment of large-scale prediction systems in high-stakes domains, e.g., face recognition, criminal justice, etc., disparity in prediction accuracy between different demographic subgroups has called for fundamental…

Machine Learning · Computer Science 2021-06-15 Jianfeng Chi , Yuan Tian , Geoffrey J. Gordon , Han Zhao

In economic program evaluation, it is common to obtain panel data in which outcomes are indicators that an individual has reached an absorbing state. For example, they may indicate whether an individual has exited a period of unemployment,…

Econometrics · Economics 2026-05-26 Ben Deaner , Hyejin Ku

Sparse latent multi-factor models have been used in many exploratory and predictive problems with high-dimensional multivariate observations. Because of concerns with identifiability, the latent factors are almost always assumed to be…

Applications · Statistics 2013-12-09 Vinicius Diniz Mayrink , Joseph Edward Lucas

Integrating datasets from different disciplines is hard because the data are often qualitatively different in meaning, scale, and reliability. When two datasets describe the same entities, many scientific questions can be phrased around…

Understanding pattern formation in crossing pedestrian flows is essential for analyzing and managing high-density crowd dynamics in urban environments. This study presents two complementary methodological approaches to detect and…

Physics and Society · Physics 2025-04-24 Piotr Nyczka , Pratik Mullick

We introduce a factor analysis model that summarizes the dependencies between observed variable groups, instead of dependencies between individual variables as standard factor analysis does. A group may correspond to one view of the same…

Machine Learning · Statistics 2014-11-19 Seppo Virtanen , Arto Klami , Suleiman A. Khan , Samuel Kaski

Generative models are invaluable in many fields of science because of their ability to capture high-dimensional and complicated distributions, such as photo-realistic images, protein structures, and connectomes. How do we evaluate the…

Difference-in-differences (diff-in-diff) is a study design that compares outcomes of two groups (treated and comparison) at two time points (pre- and post-treatment) and is widely used in evaluating new policy implementations. For instance,…

Applications · Statistics 2019-11-28 Bret Zeldow , Laura A. Hatfield

We use variation of test scores measuring closely related skills to isolate peer effects. The intuition for our identification strategy is that the difference in closely related scores eliminates factors common to the performance in either…

General Economics · Economics 2025-07-03 Guido Kuersteiner , Ingmar Prucha , Ying Zeng

Distance queries are a basic tool in data analysis. They are used for detection and localization of change for the purpose of anomaly detection, monitoring, or planning. Distance queries are particularly useful when data sets such as…

Data Structures and Algorithms · Computer Science 2015-03-20 Edith Cohen

Time series similarity measures are highly relevant in a wide range of emerging applications including training machine learning models, classification, and predictive modeling. Standard similarity measures for time series most often…

Machine Learning · Computer Science 2021-01-22 Lucas Cassiel Jacaruso

Difference-in-differences (DiD) is the most popular observational causal inference method in health policy, employed to evaluate the real-world impact of policies and programs. To estimate treatment effects, DiD relies on the "parallel…

Applications · Statistics 2024-08-09 Shuo Feng , Ishani Ganguli , Youjin Lee , John Poe , Andrew Ryan , Alyssa Bilinski

This paper is an attempt to set a justification for making use of some dicrepancy indexes, starting from the classical Maximum Likelihood definition, and adapting the corresponding basic principle of inference to situations where…

Statistics Theory · Mathematics 2021-02-24 Michel Broniatowski

Data-dependent metrics are powerful tools for learning the underlying structure of high-dimensional data. This article develops and analyzes a data-dependent metric known as diffusion state distance (DSD), which compares points using a…

Machine Learning · Statistics 2020-03-10 Lenore Cowen , Kapil Devkota , Xiaozhe Hu , James M. Murphy , Kaiyi Wu

From longitudinal biomedical studies to social networks, graphs have emerged as a powerful framework for describing evolving interactions between agents in complex systems. In such studies, after pre-processing, the data can be represented…

Applications · Statistics 2018-03-12 Claire Donnat , Susan Holmes

Statistical moments of the intensity distributions are used as molecular descriptors. They are used as a basis for defining similarity distances between two model spectra. Parameters which carry the information derived from the comparison…

Data Analysis, Statistics and Probability · Physics 2016-09-08 Dorota Bielinska-Waz , Piotr Waz , Subhash C. Basak

We propose a metric for evaluating the generalization ability of deep neural networks trained with mini-batch gradient descent. Our metric, called gradient disparity, is the $\ell_2$ norm distance between the gradient vectors of two…

Machine Learning · Computer Science 2021-07-15 Mahsa Forouzesh , Patrick Thiran

This paper explores the homogeneity of coefficients in high-dimensional regression, which extends the sparsity concept and is more general and suitable for many applications. Homogeneity arises when one expects regression coefficients…

Methodology · Statistics 2013-04-01 Tracy Ke , Jianqing Fan , Yichao Wu

Diffusion models trained on different, non-overlapping subsets of a dataset often produce strikingly similar outputs when given the same noise seed. We trace this consistency to a simple linear effect: the shared Gaussian statistics across…

Machine Learning · Computer Science 2026-02-04 Binxu Wang , Jacob Zavatone-Veth , Cengiz Pehlevan