English
Related papers

Related papers: Correcting Nonresponse Bias Using Panel Data on Da…

200 papers

The provided contents by information retrieval (IR) systems can reflect the existing societal biases and stereotypes. Such biases in retrieval results can lead to further establishing and strengthening stereotypes in society and also in the…

Information Retrieval · Computer Science 2023-01-10 Klara Krieg , Emilia Parada-Cabaleiro , Gertraud Medicus , Oleg Lesota , Markus Schedl , Navid Rekabsaz

Often, government agencies and survey organizations know the population counts or percentages for some of the variables in a survey. These may be available from auxiliary sources, for example, administrative databases or other high quality…

Methodology · Statistics 2020-11-12 Olanrewaju Akande , Gabriel Madson , D. Sunshine Hillygus , Jerome P. Reiter

Nonresponse bias is a widely prevalent problem for data on education. We develop a ten-step exemplar to guide nonresponse bias analysis (NRBA) in cross-sectional studies and apply these steps to the Early Childhood Longitudinal Study,…

Methodology · Statistics 2022-07-27 Yajuan Si , Roderick J. A. Little , Ya Mo , Nell Sedransk

Gender imbalances in work environments have been a long-standing concern. Identifying the existence of such imbalances is key to designing policies to help overcome them. In this work, we study gender trends in employment across various…

Social and Information Networks · Computer Science 2018-03-23 Karri Haranko , Emilio Zagheni , Kiran Garimella , Ingmar Weber

Recommender systems are designed to learn user preferences from observed feedback and comprise many fundamental tasks, such as rating prediction and post-click conversion rate (pCVR) prediction. However, the observed feedback usually suffer…

Information Retrieval · Computer Science 2024-02-09 Jun Wang , Haoxuan Li , Chi Zhang , Dongxu Liang , Enyun Yu , Wenwu Ou , Wenjia Wang

We present an approach to inform decisions about nonresponse follow-up sampling. The basic idea is (i) to create completed samples by imputing nonrespondents' data under various assumptions about the nonresponse mechanisms, (ii) take…

Methodology · Statistics 2022-09-16 Thais Paiva , Jerry Reiter

From scientific experiments to online A/B testing, the previously observed data often affects how future experiments are performed, which in turn affects which data will be collected. Such adaptivity introduces complex correlations between…

Machine Learning · Statistics 2018-01-03 Xinkun Nie , Xiaoying Tian , Jonathan Taylor , James Zou

We propose a way to remove the bias of a Poisson regression when the subjects are partially observed. In this paper we address this issue under certain assumptions about the missing-data generating process. We fix the total number of…

Statistics Theory · Mathematics 2014-07-08 Seyed Jalil Kazemitabar

Nonresponse is present in almost all surveys and can severely bias estimates. It is usually distinguished between unit and item nonresponse: in the former, we completely fail to have information from a unit selected in the sample, while in…

Methodology · Statistics 2015-08-25 Alina Matei , M. Giovanna Ranalli

Respondent-driven sampling (RDS) is a commonly used substitute for random sampling when studying hidden populations, such as injecting drug users or men who have sex with men, for which no sampling frame is known. The method is an extension…

Methodology · Statistics 2012-05-01 Xin Lu , Jens Malmros , Fredrik Liljeros , Tom Britton

A central goal of survey research is to collect robust and reliable data from respondents. However, despite researchers' best efforts in designing questionnaires, respondents may experience difficulty understanding questions' intent and…

Human-Computer Interaction · Computer Science 2020-11-16 Amanda Fernández-Fontelo , Pascal J. Kieslich , Felix Henninger , Frauke Kreuter , Sonja Greven

Conjoint experiments randomize multidimensional profiles, offering a powerful design for recovering structural preference parameters -- including marginal rates of substitution, willingness to pay, and the distribution of preferences across…

Methodology · Statistics 2026-05-26 Avidit Acharya , Jens Hainmueller , Yiqing Xu

Online dating platforms have gained widespread popularity as a means for individuals to seek potential romantic relationships. While recommender systems have been designed to improve the user experience in dating platforms by providing…

Information Retrieval · Computer Science 2024-02-21 Yuying Zhao , Yu Wang , Yi Zhang , Pamela Wisniewski , Charu Aggarwal , Tyler Derr

Panel data, in which multiple units are repeatedly observed over time, arise throughout science and engineering. Quantifying predictive uncertainty in such settings is challenging because conformal prediction, while distribution-free and…

Machine Learning · Statistics 2026-05-19 Daohong Tu , Kay Giesecke

In machine learning, a bias occurs whenever training sets are not representative for the test data, which results in unreliable models. The most common biases in data are arguably class imbalance and covariate shift. In this work, we aim to…

Machine Learning · Computer Science 2018-04-04 Patrick Glauner , Radu State , Petko Valtchev , Diogo Duarte

We apply a pseudo panel analysis of survey data from the years 2010 and 2017 about Americans' self-reported marital preferences and perform some formal tests on the sign and magnitude of the change in educational homophily from the…

General Economics · Economics 2025-07-24 Anna Naszodi

Massive amounts of data are the foundation of data-driven recommendation models. As an inherent nature of big data, data heterogeneity widely exists in real-world recommendation systems. It reflects the differences in the properties among…

Information Retrieval · Computer Science 2023-05-26 Zimu Wang , Jiashuo Liu , Hao Zou , Xingxuan Zhang , Yue He , Dongxu Liang , Peng Cui

Large language models (LLMs) are becoming increasingly ubiquitous in our daily lives, but numerous concerns about bias in LLMs exist. This study examines how gender-diverse populations perceive bias, accuracy, and trustworthiness in LLMs,…

Human-Computer Interaction · Computer Science 2025-07-09 Aimen Gaba , Emily Wall , Tejas Ramkumar Babu , Yuriy Brun , Kyle Hall , Cindy Xiong Bearfield

Imbalanced data, where the positive samples represent only a small proportion compared to the negative samples, makes it challenging for classification problems to balance the false positive and false negative rates. A common approach to…

Machine Learning · Statistics 2026-02-17 Pengfei Lyu , Zhengchi Ma , Linjun Zhang , Anru R. Zhang

The goal of question answering (QA) is to answer any question. However, major QA datasets have skewed distributions over gender, profession, and nationality. Despite that skew, model accuracy analysis reveals little evidence that accuracy…

Computation and Language · Computer Science 2021-09-14 Maharshi Gor , Kellie Webster , Jordan Boyd-Graber