English
Related papers

Related papers: Rethinking Factor Loading Thresholds: A Case for a…

200 papers

We propose a new active learning (AL) framework, Active Learning++, which can utilize an annotator's labels as well as its rationale. Annotators can provide their rationale for choosing a label by ranking input features based on their…

Machine Learning · Computer Science 2020-09-11 Bhavya Ghai , Q. Vera Liao , Yunfeng Zhang , Klaus Mueller

Scientific claim verification, the task of determining whether claims are entailed by scientific evidence, is fundamental to establishing discoveries in evidence while preventing misinformation. This process involves evaluating each…

Computation and Language · Computer Science 2026-04-14 Muxin Liu , Delip Rao , Grace Kim , Chris Callison-Burch

In M-open problems where no true model can be conceptualized, it is common to back off from modeling and merely seek good prediction. Even in M-complete problems, taking a predictive approach can be very useful. Stacking is a model…

Statistics Theory · Mathematics 2016-02-17 Tri Le , Bertrand Clarke

Given $n$ independent random vectors with common density $f$ on $\mathbb{R}^d$, we study the weak convergence of three empirical-measure based estimators of the convex $\lambda$-level set $L_\lambda$ of $f$, namely the excess mass set, the…

Statistics Theory · Mathematics 2020-06-04 Philippe Berthet , John H. J. Einmahl

Important tasks like record linkage and extreme classification demonstrate extreme class imbalance, with 1 minority instance to every 1 million or more majority instances. Obtaining a sufficient sample of all classes, even just to achieve…

Machine Learning · Computer Science 2021-06-03 Neil G. Marchant , Benjamin I. P. Rubinstein

A conjecture of Talagrand (2010) states that the so-called expectation and fractional expectation thresholds are always within at most some constant factor from each other. We prove for the unweighted case that this is a.a.s. true when the…

Combinatorics · Mathematics 2025-10-22 Thomas Fischer , Yury Person

A new weak measurement procedure is introduced for finite samples which yields accurate weak values that are outside the range of eigenvalues and which do not require an exponentially rare ensemble. This procedure provides a unique…

Quantum Physics · Physics 2009-11-13 Jeff Tollaksen

We propose FACTER, a fairness-aware framework for LLM-based recommendation systems that integrates conformal prediction with dynamic prompt engineering. By introducing an adaptive semantic variance threshold and a violation-triggered…

Information Retrieval · Computer Science 2025-02-06 Arya Fayyazi , Mehdi Kamal , Massoud Pedram

Instrumental variables (IVs) are extensively used to estimate treatment effects when the treatment and outcome are confounded by unmeasured confounders; however, weak IVs are often encountered in empirical studies and may cause problems.…

Methodology · Statistics 2021-10-19 Siyu Heng , Bo Zhang , Xu Han , Scott A. Lorch , Dylan S. Small

This paper proposes maximum (quasi)likelihood estimation for high dimensional factor models with regime switching in the loadings. The model parameters are estimated jointly by the EM (expectation maximization) algorithm, which in the…

Econometrics · Economics 2023-04-11 Giovanni Urga , Fa Wang

Process capability indices such as $C_{pk}$ are widely used in manufacturing quality control to support supplier qualification and product release decisions based on fixed acceptance thresholds (e.g., $C_{pk} \geq 1.33$). In practice, these…

Applications · Statistics 2026-03-13 Fei Jiang , Lei Yang

Generalized latent factor analysis not only provides a useful latent embedding approach in statistics and machine learning, but also serves as a widely used tool across various scientific fields, such as psychometrics, econometrics, and…

Methodology · Statistics 2025-08-11 Chengyu Cui , Gongjun Xu

Cited RAG evaluation often treats visible sources as a grounding signal, but a real, topically relevant citation can still under-warrant the attached wording. We study this diagnostic failure as citation laundering: a related source is…

Artificial Intelligence · Computer Science 2026-05-28 Pin Qian , Su Wang , Xiaoyuan Wang , Yihang Chen , Wenxuan Xu , Qiaolin Yu , Shuhuai Lin , Sipeng Zhang , Junxian You , Xinpeng Wei

Identifying the number of factors in a high-dimensional factor model has attracted much attention in recent years and a general solution to the problem is still lacking. A promising ratio estimator based on the singular values of the lagged…

Methodology · Statistics 2018-01-23 Zeng Li , Qinwen Wang , Jianfeng Yao

The study of model bias and variance with respect to decision boundaries is critically important in supervised classification. There is generally a tradeoff between the two, as fine-tuning of the decision boundary of a classification model…

Machine Learning · Computer Science 2020-02-25 Matthew Almeida , Wei Ding , Scott Crouter , Ping Chen

Reliable deployment of Vision-Language Models (VLMs) in radiology requires validation metrics that go beyond surface-level text similarity to ensure clinical fidelity and demographic fairness. This paper investigates a critical blind spot…

Computation and Language · Computer Science 2026-03-03 Aditya Parikh , Aasa Feragen , Sneha Das , Stella Frank

Context: Expert judgement is a common method for software effort estimations in practice today. Estimators are often shown extra obsolete requirements together with the real ones to be implemented. Only one previous study has been conducted…

Software Engineering · Computer Science 2021-03-25 Lucas Gren , Richard Berntsson Svensson

Foundation models often generate unreliable answers, while heuristic uncertainty estimators fail to fully distinguish correct from incorrect outputs, causing users to accept erroneous answers without any statistical guarantee. We address…

Artificial Intelligence · Computer Science 2026-05-27 Zhiyuan Wang , Aniri , Tianlong Chen , Yue Zhang , Heng Tao Shen , Xiaoshuang Shi , Kaidi Xu

Using instruments comprising ordered responses to items are ubiquitous for studying many constructs of interest. However, using such an item response format may lead to items with response categories infrequently endorsed or unendorsed…

Methodology · Statistics 2024-05-02 R. Noah Padgett , Grant B. Morgan , Tim Lomas

We prove a theorem justifying the regularity conditions which are needed for Path Sampling in Factor Models. We then show that the remaining ingredient, namely, MCMC for calculating the integrand at each point in the path, may be seriously…

Computation · Statistics 2013-02-22 Ritabrata Dutta , Jayanta K. Ghosh