English
Related papers

Related papers: Correcting for Nonignorable Nonresponse Bias in Or…

200 papers

Chain-of-thought reasoning enables large language models to solve multi-step tasks by framing problem solving as sequential decision problems. Outcome-based rewards, which provide feedback only on final answers, show impressive success, but…

Machine Learning · Computer Science 2025-04-15 Tarun Chitra

In this paper, we expand the Bayesian persuasion framework to account for unobserved confounding variables in sender-receiver interactions. While traditional models assume that belief updates follow Bayesian principles, real-world scenarios…

Artificial Intelligence · Computer Science 2025-08-11 Nishanth Venkatesh S. , Heeseung Bang , Andreas A. Malikopoulos

Should prediction models always deliver a prediction? In the pursuit of maximum predictive performance, critical considerations of reliability and fairness are often overshadowed, particularly when it comes to the role of uncertainty.…

Machine Learning · Computer Science 2024-10-29 Anna Sokol , Nuno Moniz , Nitesh Chawla

Nonresponse bias is a widely prevalent problem for data on education. We develop a ten-step exemplar to guide nonresponse bias analysis (NRBA) in cross-sectional studies and apply these steps to the Early Childhood Longitudinal Study,…

Methodology · Statistics 2022-07-27 Yajuan Si , Roderick J. A. Little , Ya Mo , Nell Sedransk

In content-based online platforms, use of aggregate user feedback (say, the sum of votes) is commonplace as the "gold standard" for measuring content quality. Use of vote aggregates, however, is at odds with the existing empirical…

Social and Information Networks · Computer Science 2019-10-03 Himel Dev , Karrie Karahalios , Hari Sundaram

Analysis of sample survey data often requires adjustments to account for missing data in the outcome variables of principal interest. Standard adjustment methods based on item imputation or on propensity weighting factors rely heavily on…

Methodology · Statistics 2016-03-08 Wei-Yin Loh , John Eltinge , MoonJung Cho , Yuanzhi Li

Respondent-driven sampling is a form of link-tracing network sampling, which is widely used to study hard-to-reach populations, often to estimate population proportions. Previous treatments of this process have used a with-replacement…

Methodology · Statistics 2010-06-25 Krista J. Gile

Propensity score weighting is widely used to improve the representativeness and correct the selection bias in the voluntary sample. The propensity score is often developed using a model for the sampling probability, which can be subject to…

Methodology · Statistics 2022-07-20 Hengfang Wang , Jae Kwang Kim

Estimating a causal effect from observational data can be biased if we do not control for self-selection. This selection is based on confounding variables that affect the treatment assignment and the outcome. Propensity score methods aim to…

Econometrics · Economics 2021-09-10 Daniel Jacob

In this paper we propose a strategy for administering a survey that is mindful of sensitive data and individual privacy. The survey in question seeks to estimate the population proportions of a sensitive, polychotomous variable and does not…

Statistics Theory · Mathematics 2007-06-13 Fernando Esponda

Omitted variables are one of the most important threats to the identification of causal effects. Several widely used methods assess the impact of omitted variables on empirical conclusions by comparing measures of selection on observables…

Econometrics · Economics 2026-02-05 Paul Diegert , Matthew A. Masten , Alexandre Poirier

Large language models (LLMs) can often accurately describe probability distributions using natural language, yet they still struggle to generate faithful samples from them. This mismatch limits their use in tasks requiring reliable…

Machine Learning · Computer Science 2026-04-24 Tim Z. Xiao , Johannes Zenn , Zhen Liu , Weiyang Liu , Robert Bamler , Bernhard Schölkopf

Over the past few years, question answering and information retrieval systems have become widely used. These systems attempt to find the answer of the asked questions from raw text sources. A component of these systems is Answer Selection…

Computation and Language · Computer Science 2019-11-13 Jamshid Mozafari , Mohammad Ali Nematbakhsh , Afsaneh Fatemi

It is often of interest to estimate regression functions non-parametrically. Penalized regression (PR) is one statistically-effective, well-studied solution to this problem. Unfortunately, in many cases, finding exact solutions to PR…

Methodology · Statistics 2021-12-08 Brayan Ortiz , Noah Simon

Evaluation of search engines relies on assessments of search results for selected test queries, from which we would ideally like to draw conclusions in terms of relevance of the results for general (e.g., future, unknown) users. In practice…

Information Retrieval · Computer Science 2015-11-24 Thomas Demeester , Robin Aly , Djoerd Hiemstra , Dong Nguyen , Chris Develder

Corrupted data sets containing noisy or missing observations are prevalent in various contemporary applications such as economics, finance and bioinformatics. Despite the recent methodological and algorithmic advances in high-dimensional…

Methodology · Statistics 2020-05-12 J. Wu , Z. Zheng , Y. Li , Y. Zhang

Multivariate matched proportions (MMP) data appears in a variety of contexts including post-market surveillance of adverse events in pharmaceuticals, disease classification, and agreement between care providers. It consists of multiple sets…

Methodology · Statistics 2023-05-08 Mark J. Meyer , Haobo Cheng , Katherine Hobbs Knutson

Nonresponse in panel studies can lead to a substantial loss in data quality due to its potential to introduce bias and distort survey estimates. Recent work investigates the usage of machine learning to predict nonresponse in advance, such…

Methodology · Statistics 2019-11-05 Christoph Kern , Bernd Weiss , Jan-Philipp Kolb

Biomedical studies have a common interest in assessing relationships between multiple related health outcomes and high-dimensional predictors. For example, in reproductive epidemiology, one may collect pregnancy outcomes such as length of…

Applications · Statistics 2011-11-24 Bin Zhu , David B. Dunson , Allison E. Ashley-Koch

Large Language Models (LLMs) are increasingly used to simulate population responses, a method known as ``Silicon Sampling''. However, responses to socially sensitive questions frequently exhibit Social Desirability Bias (SDB), diverging…

Computation and Language · Computer Science 2025-12-30 Sashank Chapala , Maksym Mironov , Songgaojun Deng
‹ Prev 1 4 5 6 7 8 10 Next ›