English
Related papers

Related papers: Know your population and know your model: Using mo…

200 papers

Cognitive biases are widespread in humans and animals alike, and can sometimes be reinforced by social interactions. One prime bias in judgment and decision-making is the human tendency to underestimate large quantities. Previous research…

Physics and Society · Physics 2022-01-12 Bertrand Jayles , Clément Sire , Ralf H. J. M Kurvers

Causal inference in a program evaluation setting faces the problem of external validity when the treatment effect in the target population is different from the treatment effect identified from the population of which the sample is…

Methodology · Statistics 2021-12-23 Kyungchul Song , Zhengfei Yu

Machine learning is often viewed as an inherently value-neutral process: statistical tendencies in the training inputs are "simply" used to generalize to new examples. However when models impact social systems such as interactions between…

Computers and Society · Computer Science 2019-08-21 Ben Hutchinson , KJ Pittl , Margaret Mitchell

Investigators often use multi-source data (e.g., multi-center trials, meta-analyses of randomized trials, pooled analyses of observational cohorts) to learn about the effects of interventions in subgroups of some well-defined target…

Methodology · Statistics 2024-02-06 Guanbo Wang , Alexander Levis , Jon Steingrimsson , Issa Dahabreh

Estimation of social influence in networks can be substantially biased in observational studies due to homophily and network correlation in exposure to exogenous events. Randomized experiments, in which the researcher intervenes in the…

Social and Information Networks · Computer Science 2017-09-28 Sean J. Taylor , Dean Eckles

Differentially private (DP) mechanisms have been deployed in a variety of high-impact social settings (perhaps most notably by the U.S. Census). Since all DP mechanisms involve adding noise to results of statistical queries, they are…

Cryptography and Security · Computer Science 2023-12-20 Lucas Rosenblatt , Julia Stoyanovich , Christopher Musco

In using observed data to make inferences about a population quantity, it is commonly assumed that the sampling distribution from which the data were drawn belongs to a given parametric family of distributions, or at least, a given finite…

Methodology · Statistics 2024-10-21 Russell J. Bowater

It has become apparent that models that have been applied widely in economics, including Machine Learning techniques and Data Mining methods, should take into consideration principles that derive from the theories of Personality Psychology…

Machine Learning · Computer Science 2013-07-09 Alexandros Ladas , Uwe Aickelin , Jon Garibaldi , Eamonn Ferguson

Diffusion probabilistic models have been successfully used to generate data from noise. However, most diffusion models are computationally expensive and difficult to interpret with a lack of theoretical justification. Random feature models…

Machine Learning · Statistics 2025-08-11 Esha Saha , Giang Tran

Data augmentation is an important technique in training deep neural networks as it enhances their ability to generalize and remain robust. While data augmentation is commonly used to expand the sample size and act as a consistency…

Machine Learning · Computer Science 2025-02-18 Xiliang Yang , Shenyang Deng , Shicong Liu , Yuanchi Suo , Wing. W. Y NG , Jianjun Zhang

When studying policy interventions, researchers often pursue two goals: i) identifying for whom the program has the largest effects (heterogeneity) and ii) determining whether those patterns of treatment effects have predictive power across…

Econometrics · Economics 2025-07-28 Emily Breza , Arun G. Chandrasekhar , Davide Viviano

We consider settings where the observations are drawn from a zero-mean multivariate (real or complex) normal distribution with the population covariance matrix having eigenvalues of arbitrary multiplicity. We assume that the eigenvectors of…

Statistics Theory · Mathematics 2009-01-22 N. Raj Rao , James A. Mingo , Roland Speicher , Alan Edelman

Although randomized controlled trials have long been regarded as the ``gold standard'' for evaluating treatment effects, there is no natural prevention from post-treatment events. For example, non-compliance makes the actual treatment…

Methodology · Statistics 2025-04-25 Qinqing Liu , Xiang Peng , Tao Zhang , Yuhao Deng

Restricting randomization in the design of experiments (e.g., using blocking/stratification, pair-wise matching, or rerandomization) can improve the treatment-control balance on important covariates and therefore improve the estimation of…

Econometrics · Economics 2020-11-02 Brian Quistorff , Gentry Johnson

Statistical Relational Learning (SRL) methods have shown that classification accuracy can be improved by integrating relations between samples. Techniques such as iterative classification or relaxation labeling achieve this by propagating…

Information Retrieval · Computer Science 2017-02-13 Immanuel Bayer , Uwe Nagel , Steffen Rendle

Lack of repeatability and generalisability are two significant threats to continuing scientific development in Natural Language Processing. Language models and learning methods are so complex that scientific conference papers no longer…

Computation and Language · Computer Science 2018-08-07 Andrew Moore , Paul Rayson

This paper considers the problem of estimation in the generalized semiparametric model for longitudinal data when the number of parameters diverges with the sample size. A penalization type of generalized estimating equation method is…

Methodology · Statistics 2020-06-09 M. Taavoni , M. Arashi

There is wide agreement on the importance of implementation data from randomized effectiveness studies in behavioral science; however, there are few methods available to incorporate these data into causal models, especially when they are…

Methodology · Statistics 2024-05-17 Sooyong Lee , Adam C Sales , Hyeon-Ah Kang , Tiffany A. Whittaker

This paper is about how we study statistical methods. As an example, it uses the random regressions model, in which the intercept and slope of cluster-specific regression lines are modeled as a bivariate random effect. Maximizing this…

Other Statistics · Statistics 2019-05-22 James S. Hodges

The Big Data revolution is challenging the state-of-the-art statistical and econometric techniques not only for the computational burden connected with the high volume and speed which data are generated, but even more for the variety of…

Methodology · Statistics 2024-10-25 Giuseppe Arbia , Vincenzo Nardelli