English
Related papers

Related papers: Correcting Sample Selection Bias in PISA Rankings

200 papers

A learned generative model often produces biased statistics relative to the underlying data distribution. A standard technique to correct this bias is importance sampling, where samples from the model are weighted by the likelihood ratio…

Machine Learning · Statistics 2019-11-05 Aditya Grover , Jiaming Song , Alekh Agarwal , Kenneth Tran , Ashish Kapoor , Eric Horvitz , Stefano Ermon

Lexicase selection is a parent selection method that considers training cases individually, rather than in aggregate, when performing parent selection. Whereas previous work has demonstrated the ability of lexicase selection to solve…

Neural and Evolutionary Computing · Computer Science 2018-05-01 William La Cava , Thomas Helmuth , Lee Spector , Jason H. Moore

The paper opens with an overview of the discussion of international comparisons (including goals) in mathematics education. Afterwards, the two most important recent international studies, the PISA Study and TIMSS-Repeat, are described.…

History and Overview · Mathematics 2007-05-23 Gabriele Kaiser , Frederick K. S. Leung , Thomas Romberg , Ivan Yaschenko

Clinical study populations often differ meaningfully from the broader populations to which results are intended to generalize. Weighting methods such as inverse probability of sampling weights (IPSW) reweight study participants to resemble…

Methodology · Statistics 2025-12-02 William Stewart , Carly L. Brantner , Elizabeth A. Stuart , Laine Thomas

Recent empirical work has shown that human children are adept at learning and reasoning with probabilities. Here, we model a recent experiment investigating the development of school-age children's non-symbolic probability reasoning ability…

Neurons and Cognition · Quantitative Biology 2023-05-09 Zilong Wang , Thomas R. Shultz , Ardvan S. Nobandegani

Computerized Adaptive Testing (CAT) is a widely used technology for evaluating learners' proficiency in online education platforms. By leveraging prior estimates of proficiency to select questions and updating the estimates iteratively…

Information Retrieval · Computer Science 2025-12-24 Mi Tian , Kun Zhang , Fei Liu , Jinglong Li , Yuxin Liao , Chenxi Bai , Zhengtao Tan , Le Wu , Richang Hong

A common issue for classification in scientific research and industry is the existence of imbalanced classes. When sample sizes of different classes are imbalanced in training data, naively implementing a classification method often leads…

Methodology · Statistics 2021-07-02 Yang Feng , Min Zhou , Xin Tong

Panel studies typically suffer from attrition, which reduces sample size and can result in biased inferences. It is impossible to know whether or not the attrition causes bias from the observed panel data alone. Refreshment samples - new,…

Methodology · Statistics 2013-06-13 Yiting Deng , D. Sunshine Hillygus , Jerome P. Reiter , Yajuan Si , Siyu Zheng

In predictive tasks, real-world datasets often present different degrees of imbalanced (i.e., long-tailed or skewed) distributions. While the majority (the head) classes have sufficient samples, the minority (the tail) classes can be…

Machine Learning · Computer Science 2021-09-14 Chongsheng Zhang , Paolo Soda , Jingjun Bi , Gaojuan Fan , George Almpanidis , Salvador Garcia

I have three goals in this article: (1) To show the enormous potential of bootstrapping and permutation tests to help students understand statistical concepts including sampling distributions, standard errors, bias, confidence intervals,…

Other Statistics · Statistics 2014-11-20 Tim Hesterberg

In classification problems, sampling bias between training data and testing data is critical to the ranking performance of classification scores. Such bias can be both unintentionally introduced by data collection and intentionally…

Methodology · Statistics 2017-11-02 Chandler Zuo

This article is devoted to the problem of predicting the value taken by a random permutation $\Sigma$, describing the preferences of an individual over a set of numbered items $\{1,\; \ldots,\; n\}$ say, based on the observation of an…

Statistics Theory · Mathematics 2017-12-20 Stephan Clémençon , Anna Korba , Eric Sibony

As deep learning continues to be driven by ever-larger datasets, understanding which examples are most important for generalization has become a critical question. While progress in data selection continues, emerging applications require…

Machine Learning · Computer Science 2025-07-02 Mustafa Burak Gurbuz , Xingyu Zheng , Constantine Dovrolis

Researchers increasingly have access to two types of data: (i) large observational datasets where treatment (e.g., class size) is not randomized but several primary outcomes (e.g., graduation rates) and secondary outcomes (e.g., test…

Methodology · Statistics 2025-05-29 Susan Athey , Raj Chetty , Guido Imbens

Recent advances in unbiased learning to rank (LTR) count on Inverse Propensity Scoring (IPS) to eliminate bias in implicit feedback. Though theoretically sound in correcting the bias introduced by treating clicked documents as relevant, IPS…

Information Retrieval · Computer Science 2021-11-16 Nan Wang , Zhen Qin , Xuanhui Wang , Hongning Wang

Network sampling is used around the world for surveys of vulnerable, hard-to-reach populations including people at risk for HIV, opioid misuse, and emerging epidemics. The sampling methods include tracing social links to add new people to…

Methodology · Statistics 2020-02-05 Steve Thompson

Selection bias affects Mendelian randomization investigations when selection into the study sample depends on a collider between the genetic variant and confounders of the risk factor-outcome association. However, the relative importance of…

Applications · Statistics 2018-03-13 Apostolos Gkatzionis , Stephen Burgess

Many partial identification problems can be characterized by the optimal value of a function over a set where both the function and set need to be estimated by empirical data. Despite some progress for convex problems, statistical inference…

Methodology · Statistics 2022-08-31 Matthew Tudball , Rachael Hughes , Kate Tilling , Jack Bowden , Qingyuan Zhao

Despite the increasing use of citation-based metrics for research evaluation purposes, we do not know yet which metrics best deliver on their promise to gauge the significance of a scientific paper or a patent. We assess 17 network-based…

Social and Information Networks · Computer Science 2020-07-10 Shuqi Xu , Manuel Sebastian Mariani , Linyuan Lü , Matúš Medo

It has long been noticed that the efficacy observed in small early phase studies is generally better than that observed in later larger studies. Historically, the inflation of the efficacy results from early proof-of-concept studies is…

Methodology · Statistics 2020-06-11 Yongming Qu , Yu Du , Ying Zhang , Lei Shen