English
Related papers

Related papers: Know your population and know your model: Using mo…

200 papers

Personality and demographics are important variables in social sciences, while in NLP they can aid in interpretability and removal of societal biases. However, datasets with both personality and demographic labels are scarce. To address…

Computation and Language · Computer Science 2021-06-09 Matej Gjurković , Mladen Karan , Iva Vukojević , Mihaela Bošnjak , Jan Šnajder

Real-world generalization, e.g., deciding to approach a never-seen-before animal, relies on contextual information as well as previous experiences. Such a seemingly easy behavioral choice requires the interplay of multiple neural…

Neurons and Cognition · Quantitative Biology 2022-01-17 Peer Herholz , Eddy Fortier , Mariya Toneva , Nicolas Farrugia , Leila Wehbe , Valentina Borghesani

Neurons can code for multiple variables simultaneously and neuroscientists are often interested in classifying neurons based on their receptive field properties. Statistical models provide powerful tools for determining the factors…

Neurons and Cognition · Quantitative Biology 2022-10-28 Mehrad Sarmashghi , Shantanu P. Jadhav , Uri T. Eden

Modal regression, a widely used regression protocol, has been extensively investigated in statistical and machine learning communities due to its robustness to outliers and heavy-tailed noises. Understanding modal regression's theoretical…

Machine Learning · Statistics 2022-03-15 Tielang Gong , Yuxin Dong , Hong Chen , Bo Dong , Wei Feng , Chen Li

Several techniques exist to assess and reduce nonresponse bias, including propensity models, calibration methods, or post-stratification. These approaches can only be applied after the data collection, and assume reliable information…

Methodology · Statistics 2020-05-26 Blanka Szeitl , Tamás Rudas

The principle of maximum entropy provides a useful method for inferring statistical mechanics models from observations in correlated systems, and is widely used in a variety of fields where accurate data are available. While the assumptions…

Neurons and Cognition · Quantitative Biology 2017-06-02 Ulisse Ferrari , Tomoyuki Obuchi , Thierry Mora

Post-treatment variables often complicate causal inference. They appear in many scientific problems, including noncompliance, truncation by death, mediation, and surrogate endpoint evaluation. Principal stratification is a strategy to…

Methodology · Statistics 2024-04-04 Sizhu Lu , Zhichao Jiang , Peng Ding

Respondent-driven sampling is a form of link-tracing network sampling, which is widely used to study hard-to-reach populations, often to estimate population proportions. Previous treatments of this process have used a with-replacement…

Methodology · Statistics 2010-06-25 Krista J. Gile

Recent advancements in improving the reasoning capabilities of Large Language Models have underscored the efficacy of Process Reward Models (PRMs) in addressing intermediate errors through structured feedback mechanisms. This study analyzes…

Computation and Language · Computer Science 2025-06-03 Zhengyu Chen , Yudong Wang , Teng Xiao , Ruochen Zhou , Xuesheng Yang , Wei Wang , Zhifang Sui , Jingang Wang

Consider a panel data setting where repeated observations on individuals are available. Often it is reasonable to assume that there exist groups of individuals that share similar effects of observed characteristics, but the grouping is…

Methodology · Statistics 2024-02-09 Lu Yu , Jiaying Gu , Stanislav Volgushev

Mendelian Randomization (MR) is a prominent observational epidemiological research method designed to address unobserved confounding when estimating causal effects. However, core assumptions -- particularly the independence between…

Machine Learning · Computer Science 2026-02-24 Shimeng Huang , Matthew Robinson , Francesco Locatello

Principal stratification (PS) is a commonly used approach for understanding the mechanisms through which a treatment affects an outcome. The goal of this work is to extend the PS framework to studies with continuous treatments, which…

Methodology · Statistics 2025-05-20 Joseph Antonelli , Minxuan Wu , Fabrizia Mealli , Brenden Beck , Alessandra Mattei

Randomness in scientific estimation is generally assumed to arise from unmeasured or uncontrolled factors. However, when combining subjective probability estimates, heterogeneity stemming from people's cognitive or information diversity is…

Methodology · Statistics 2015-05-28 Ville Satopää , Robin Pemantle , Lyle Ungar

Respondent-Driven Sampling (RDS) employs a variant of a link-tracing network sampling strategy to collect data from hard-to-reach populations. By tracing the links in the underlying social network, the process exploits the social structure…

Applications · Statistics 2009-04-14 Krista J. Gile , Mark S. Handcock

The aim of this paper is twofold. First, three theoretical principles are formalized: randomization, overrepresentation and restriction. We develop these principles and give a rationale for their use in choosing the sampling design in a…

Methodology · Statistics 2016-12-16 Yves Tillé , Matthieu Wilhelm

There is a growing need for flexible general frameworks that integrate individual-level data with external summary information for improved statistical inference. External information relevant for a risk prediction model may come in…

Methodology · Statistics 2023-04-11 Tian Gu , Jeremy M. G. Taylor , Bhramar Mukherjee

Prediction polling is an increasingly popular form of crowdsourcing in which multiple participants estimate the probability or magnitude of some future event. These estimates are then aggregated into a single forecast. Historically,…

Methodology · Statistics 2016-04-25 Ville A. Satopää , Shane T. Jensen , Robin Pemantle , Lyle H. Ungar

Statistical NLP systems are frequently evaluated and compared on the basis of their performances on a single split of training and test data. Results obtained using a single split are, however, subject to sampling noise. In this paper we…

Computation and Language · Computer Science 2007-05-23 Yuval Krymolowski

In real word applications, data generating process for training a machine learning model often differs from what the model encounters in the test stage. Understanding how and whether machine learning models generalize under such…

Machine Learning · Statistics 2022-02-08 Abdulkadir Canatar , Blake Bordelon , Cengiz Pehlevan

In the November 2016 U.S. presidential election, many state level public opinion polls, particularly in the Upper Midwest, incorrectly predicted the winning candidate. One leading explanation for this polling miss is that the precipitous…

Methodology · Statistics 2021-11-15 Eli Ben-Michael , Avi Feller , Erin Hartman
‹ Prev 1 4 5 6 7 8 10 Next ›