English
Related papers

Related papers: Integration of Individual Participant and Aggregat…

200 papers

Dataset distillation condenses large datasets into synthetic subsets, achieving performance comparable to training on the full dataset while substantially reducing storage and computation costs. Most existing dataset distillation methods…

Machine Learning · Computer Science 2025-11-03 Qiyan Deng , Changqian Zheng , Lianpeng Qiao , Yuping Wang , Chengliang Chai , Lei Cao

This paper proposes a cohort-anchored framework for robust inference in event studies with staggered adoption, building on Rambachan and Roth (2023). Robust inference based on event-study coefficients aggregated across cohorts can be…

Econometrics · Economics 2025-09-30 Ziyi Liu

The aim of survey statistics is to produce estimates with a minimal bias and a corresponding acceptable variance given a specific budget, preferable with a minor response burden for the participants. In recent years, considerable efforts…

Methodology · Statistics 2026-04-02 Martin Hyllienmark , Gustaf Strandell

Data augmentation (DA) is a crucial technique for enhancing the sample efficiency of visual reinforcement learning (RL) algorithms. Notably, employing simple observation transformations alone can yield outstanding performance without extra…

Machine Learning · Computer Science 2023-10-30 Guozheng Ma , Linrui Zhang , Haoyu Wang , Lu Li , Zilin Wang , Zhen Wang , Li Shen , Xueqian Wang , Dacheng Tao

Varying domains and biased datasets can lead to differences between the training and the target distributions, known as covariate shift. Current approaches for alleviating this often rely on estimating the ratio of training and target…

Machine Learning · Statistics 2020-10-27 Bijan Mazaheri , Siddharth Jain , Jehoshua Bruck

We discuss how to handle matching-adjusted indirect comparison (MAIC) from a data analyst's perspective. We introduce several multivariate data analysis methods to assess the appropriateness of MAIC for a given data set. These methods focus…

Applications · Statistics 2022-03-18 Ekkehard Glimm , Lillian Yau

Aggregated Relational Data (ARD) contain summary information about individual social networks and are widely used to estimate social network characteristics and the size of populations of interest. Although a variety of ARD estimators…

Methodology · Statistics 2026-01-27 Ian Laga , Benjamin Vogel , Jieyun Wang , Anna Smith , Owen Ward

Integrative analysis of datasets generated by multiple cohorts is a widely-used approach for increasing sample size, precision of population estimators, and generalizability of analysis results in epidemiological studies. However, often…

Data augmentation (DA) is indispensable in modern machine learning and deep neural networks. The basic idea of DA is to construct new training data to improve the model's generalization by adding slightly disturbed versions of existing data…

Machine Learning · Computer Science 2024-06-05 Chengtai Cao , Fan Zhou , Yurou Dai , Jianping Wang , Kunpeng Zhang

Data fusion describes the method of combining data from (at least) two initially independent data sources to allow for joint analysis of variables which are not jointly observed. The fundamental idea is to base inference on identifying…

Methodology · Statistics 2020-12-02 Florian Meinfelder , Jannik Schaller

Big data presents potential but unresolved value as a source for analysis and inference. However,selection bias, present in many of these datasets, needs to be accounted for so that appropriate inferences can be made on the target…

Methodology · Statistics 2025-01-09 Lyndon Ang , Robert Clark , Bronwyn Loong , Anders Holmberg

Some classical uncertainty quantification problems require the estimation of multiple expectations. Estimating all of them accurately is crucial and can have a major impact on the analysis to perform, and standard existing Monte Carlo…

Methodology · Statistics 2022-12-02 Julien Demange-Chryst , François Bachoc , Jérôme Morio

We propose a distributed method for simultaneous inference for datasets with sample size much larger than the number of covariates, i.e., N >> p, in the generalized linear models framework. When such datasets are too big to be analyzed…

Methodology · Statistics 2020-07-23 Lu Tang , Ling Zhou , Peter X. -K. Song

Describing statistical dependencies is foundational to empirical scientific research. For uncovering intricate and possibly non-linear dependencies between a single target variable and several source variables within a system, a principled…

Information Theory · Computer Science 2024-03-28 David A. Ehrlich , Kyle Schick-Poland , Abdullah Makkeh , Felix Lanfermann , Patricia Wollstadt , Michael Wibral

In several social choice problems, agents collectively make decisions over the allocation of multiple divisible and heterogeneous resources with capacity constraints to maximize utilitarian social welfare. The agents are constrained through…

Optimization and Control · Mathematics 2023-11-02 Syed Eqbal Alam , Fabian Wirth , Jia Yuan Yu , Robert Shorten

The subpopulationtion shift, characterized by a disparity in subpopulation distributibetween theween the training and target datasets, can significantly degrade the performance of machine learning models. Current solutions to subpopulation…

We consider the privacy amplification properties of a sampling scheme in which a user's data is used in $k$ steps chosen randomly and uniformly from a sequence (or set) of $t$ steps. This sampling scheme has been recently applied in the…

Machine Learning · Computer Science 2026-02-20 Vitaly Feldman , Moshe Shenfeld

The Integrated Probabilistic Data Association Filter (IPDAF) is a target tracking algorithm based on the Probabilistic Data Association Filter that calculates a statistical measure that indicates if an estimated representation of the target…

Robotics · Computer Science 2022-10-18 Mark E. Petersen , Randal W. Beard

Applied analysts often use the differences-in-differences (DID) method to estimate the causal effect of policy interventions with observational data. The method is widely used, as the required before and after comparison of a treated and…

Applications · Statistics 2019-02-04 Luke J. Keele , Dylan S. Small , Jesse Y. Hsu , Colin B. Fogarty

In this paper, we develop a multiply robust inference procedure of the average treatment effect (ATE) for data with high-dimensional covariates. We consider the case where it is difficult to correctly specify a single parametric model for…

Methodology · Statistics 2025-09-03 Xintao Xia , Yumou Qiu