中文
相关论文

相关论文: Multilevel calibration weighting for survey data

200 篇论文

The alignment of large language models (LLMs) with human values increasingly relies on using other LLMs as automated judges, or ``autoraters''. However, their reliability is limited by a foundational issue: they are trained on discrete…

Generalization to new samples is a fundamental rationale for statistical modeling. For this purpose, model validation is particularly important, but recent work in survey inference has suggested that simple aggregation of individual…

统计方法学 · 统计学 2024-04-15 Lauren Kennedy , Aki Vehtari , Andrew Gelman

Modern surveys with large sample sizes and growing mixed-type questionnaires require robust and scalable analysis methods. In this work, we consider recovering a mixed dataframe matrix, obtained by complex survey sampling, with entries…

统计方法学 · 统计学 2024-02-07 Xiaojun Mao , Hengfang Wang , Zhonglei Wang , Shu Yang

Multi-calibration is a powerful and evolving concept originating in the field of algorithmic fairness. For a predictor $f$ that estimates the outcome $y$ given covariates $x$, and for a function class $\mathcal{C}$, multi-calibration…

机器学习 · 计算机科学 2023-03-09 Zhun Deng , Cynthia Dwork , Linjun Zhang

The propensity score is widely used for causal inference in observational studies, but common parametric estimators can produce biased and inefficient effect estimates when model assumptions are violated. Nonparametric approaches reduce…

统计方法学 · 统计学 2026-04-09 Maosen Peng , Yan Li , Chong Wu , Liang Li

When facing uncertainty, decision-makers want predictions they can trust. A machine learning provider can convey confidence to decision-makers by guaranteeing their predictions are distribution calibrated -- amongst the inputs that receive…

机器学习 · 统计学 2021-07-14 Shengjia Zhao , Michael P. Kim , Roshni Sahoo , Tengyu Ma , Stefano Ermon

Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) by encouraging their generations to have high rewards, using a reward model trained on human preferences. To prevent the forgetting of pre-trained…

When estimating a regression model, we might have data where some labels are missing, or our data might be biased by a selection mechanism. When the response or selection mechanism is ignorable (i.e., independent of the response variable…

统计理论 · 数学 2023-08-22 Philip Boeken , Noud de Kroon , Mathijs de Jong , Joris M. Mooij , Onno Zoeter

The general principles of Bayesian data analysis imply that models for survey responses should be constructed conditional on all variables that affect the probability of inclusion and nonresponse, which are also the variables used in survey…

统计方法学 · 统计学 2007-11-06 Andrew Gelman

Large Reasoning Models have demonstrated remarkable performance with the advancement of test-time scaling techniques, which enhances prediction accuracy by generating multiple candidate responses and selecting the most reliable answer.…

机器学习 · 计算机科学 2026-03-05 Xizhong Yang , Haotian Zhang , Huiming Wang , Mofei Song

Recalibration of binary probabilistic classifiers to a target prior probability is an important task in areas like credit risk management. However, recalibration of a classifier learned on a training dataset to a target on a test dataset in…

机器学习 · 计算机科学 2026-02-02 Dirk Tasche

In many machine learning applications, it is important for the model to provide confidence scores that accurately capture its prediction uncertainty. Although modern learning methods have achieved great success in predictive accuracy,…

机器学习 · 计算机科学 2022-07-12 Linjun Zhang , Zhun Deng , Kenji Kawaguchi , James Zou

After the revelation that neural networks tend to produce overconfident predictions, the problem of calibration, which aims to align confidence with accuracy to enhance the reliability of predictions, has gained significant importance.…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Gyusang Cho , Chan-Hyun Youn

A suitable scalar metric can help measure multi-calibration, defined as follows. When the expected values of observed responses are equal to corresponding predicted probabilities, the probabilistic predictions are known as "perfectly…

统计方法学 · 统计学 2026-04-17 Ido Guy , Daniel Haimovich , Fridolin Linder , Nastaran Okati , Lorenzo Perini , Niek Tax , Mark Tygert

Fitting mixed models to complex survey data is a challenging problem. Most methods in the literature, including the most widely used one, require a close relationship between the model structure and the survey design. In this paper we…

统计方法学 · 统计学 2023-11-23 Thomas Lumley , Xudong Huang

Propensity score weighting is a common method for estimating treatment effects with survey data. The method is applied to minimize confounding using measured covariates that are often different between individuals in treatment and control.…

统计方法学 · 统计学 2026-02-06 Yukang Zeng , Fan Li , Guangyu Tong

Large reasoning models (LRMs) show strong capabilities in complex reasoning, yet their marginal gains on evidence-dependent factual questions are limited. We find this limitation is partially attributable to a reasoning-answer hit gap,…

计算与语言 · 计算机科学 2026-01-06 Xinming Wang , Jian Xu , Bin Yu , Sheng Lian , Hongzhu Yi , Yi Chen , Yingjian Zhu , Boran Wang , Hongming Yang , Han Hu , Xu-Yao Zhang , Cheng-Lin Liu

Calibration weighting is a fundamental technique in survey sampling and data integration for incorporating auxiliary information and improving efficiency of estimators. Classical calibration methods are typically formulated through distance…

统计方法学 · 统计学 2026-03-24 Jae Kwang Kim , Yonghyun Kwon , Yumou Qiu

Multi-stage ranking pipelines have become widely used strategies in modern recommender systems, where the final stage aims to return a ranked list of items that balances a number of requirements such as user preference, diversity, novelty…

信息检索 · 计算机科学 2023-07-19 Sirui Chen , Yuan Wang , Zijing Wen , Zhiyu Li , Changshuo Zhang , Xiao Zhang , Quan Lin , Cheng Zhu , Jun Xu

Aligning large language models (LLMs) with human preferences through reinforcement learning (RLHF) can lead to reward hacking, where LLMs exploit failures in the reward model (RM) to achieve seemingly high rewards without meeting the…

‹ 上一页 1 8 9 10 下一页 ›