English
Related papers

Related papers: Using Survey Data to Obtain More Representative Si…

200 papers

Before embarking on data collection, researchers typically compute how many individual observations they should do. This is vital for doing studies with sufficient statistical power, and often a cornerstone in study pre-registrations and…

Methodology · Statistics 2023-09-06 Edwin S Dalmaijer

With the increasing pervasiveness of algorithms across industry and government, a growing body of work has grappled with how to understand their societal impact and ethical implications. Various methods have been used at different stages of…

Computers and Society · Computer Science 2022-07-21 Julia Barnett , Nicholas Diakopoulos

Data describing human interactions often suffer from incomplete sampling of the underlying population. As a consequence, the study of contagion processes using data-driven models can lead to a severe underestimation of the epidemic risk.…

Physics and Society · Physics 2015-11-19 Mathieu Génois , Christian L. Vestergaard , Ciro Cattuto , Alain Barrat

Population size estimates for hidden and hard-to-reach populations are particularly important when members are known to suffer from disproportion health issues or to pose health risks to the larger ambient population in which they are…

Social and Information Networks · Computer Science 2018-07-04 Bilal Khan , Hsuan-Wei Lee , Ian Fellows , Kirk Dombrowski

Human decision making can be challenging to predict because decisions are affected by a number of complex factors. Adding to this complexity, decision-making processes can differ considerably between individuals, and methods aimed at…

In this paper, we present a new way of matching in observational studies that overcomes three limitations of existing matching approaches. First, it directly balances covariates with multi-valued treatments without requiring the generalized…

Applications · Statistics 2019-07-11 Magdalena Bennett , Juan Pablo Vielma , Jose R. Zubizarreta

Big Data often presents as massive non-probability samples. Not only is the selection mechanism often unknown, but larger data volume amplifies the relative contribution of selection bias to total error. Existing bias adjustment approaches…

Methodology · Statistics 2022-03-29 Ali Rafei , Carol A. C. Flannagan , Brady T. West , Michael R. Elliott

Selective recruitment designs preferentially recruit individuals that are estimated to be statistically informative onto a clinical trial. Individuals that are expected to contribute less information have a lower probability of recruitment.…

Statistics Theory · Mathematics 2017-05-31 James E. Barrett

Causal inference in a program evaluation setting faces the problem of external validity when the treatment effect in the target population is different from the treatment effect identified from the population of which the sample is…

Methodology · Statistics 2021-12-23 Kyungchul Song , Zhengfei Yu

In many randomized trials, outcomes such as essays or open-ended responses must be manually scored as a preliminary step to impact analysis, a process that is costly and limiting. Model-assisted estimation offers a way to combine surrogate…

Methodology · Statistics 2026-02-16 Reagan Mozer , Nicole E. Pashley , Luke Miratrix

Crowd movement guidance has been a fascinating problem in various fields, such as easing traffic congestion in unusual events and evacuating people from an emergency-affected area. To grab the reins of crowds, there has been considerable…

Machine Learning · Computer Science 2021-07-20 Koh Takeuchi , Ryo Nishida , Hisashi Kashima , Masaki Onishi

We propose a statistical model to estimate population proportions under the survey variable cause model (Groves 2006), the setting in which the characteristic measured by the survey has a direct causal effect on survey participation. For…

Applications · Statistics 2025-06-19 Jonathan Auerbach

One way to investigate the precision of estimates likely to result from planned experiments and planned epidemiological studies is to simulate a large number of possible outcomes and analyse the sets of possible results. This appears to be…

Computation · Statistics 2013-06-28 G. K. Robinson , L. M. Ryan

In the big data era researchers face a series of problems. Even standard approaches/methodologies, like linear regression, can be difficult or problematic with huge volumes of data. Traditional approaches for regression in big datasets may…

Methodology · Statistics 2024-11-13 Vasilis Chasiotis , Dimitris Karlis

When random effects are correlated with sample design variables, the usual approach of employing individual survey weights (constructed to be inversely proportional to the unit survey inclusion probabilities) to form a pseudo-likelihood no…

Methodology · Statistics 2021-08-26 Terrance D. Savitsky , Matthew R. Williams

Non-probability sampling, for example in the form of online panels, has become a fast and cheap method to collect data. While reliable inference tools are available for classical probability samples, non-probability samples can yield…

Methodology · Statistics 2022-04-05 Gerhard Tutz

Availability, collection and access to quantitative data, as well as its limitations, often make qualitative data the resource upon which development programs heavily rely. Both traditional interview data and social media analysis can…

Computation and Language · Computer Science 2017-09-19 Philipp Broniecki , Anna Hanchar , Slava J. Mikhaylov

Randomized experiments on social networks pose statistical challenges, due to the possibility of interference between units. We propose new methods for estimating attributable treatment effects in such settings. The methods do not require…

Methodology · Statistics 2015-10-13 David S. Choi

Percentiles are statistics pointing to the standing of a paper's citation impact relative to other papers in a given citation distribution. Percentile Ranks (PRs) often play an important role in evaluating the impact of scholars,…

Digital Libraries · Computer Science 2020-05-04 Lutz Bornmann , Richard Williams

Data selection is of great significance in pre-training large language models, given the variation in quality within the large-scale available training corpora. To achieve this, researchers are currently investigating the use of data…

Artificial Intelligence · Computer Science 2024-10-08 Chi Zhang , Huaping Zhong , Kuan Zhang , Chengliang Chai , Rui Wang , Xinlin Zhuang , Tianyi Bai , Jiantao Qiu , Lei Cao , Ju Fan , Ye Yuan , Guoren Wang , Conghui He