English
Related papers

Related papers: Improved Tests for Mediation

200 papers

Test-time scaling via solution sampling and aggregation has become a key paradigm for improving the reasoning performance of Large Language Models (LLMs). While reward model selection is commonly employed in this approach, it often fails to…

Machine Learning · Computer Science 2025-09-30 Zhicheng Yang , Zhijiang Guo , Yinya Huang , Yongxin Wang , Yiwei Wang , Xiaodan Liang , Jing Tang

Large language models (LLMs) have been proposed as scalable tools to address the gap between the importance of individualized written feedback and the practical challenges of providing it at scale. However, concerns persist regarding the…

Other Statistics · Statistics 2025-11-12 Niklas Ippisch , Markus Herklotz , Anna-Carolina Haensch , Carsten Schwemmer

Safety-critical robot systems need thorough testing to expose design flaws and software bugs which could endanger humans. Testing in simulation is becoming increasingly popular, as it can be applied early in the development process and does…

The classical D'Alembert's Ratio Test is a powerful test that we learn from calculus to determine convergence for a series of positive terms. Its range of applicability and ease of computation makes this test extremely appealing. However,…

Classical Analysis and ODEs · Mathematics 2021-09-16 Edward Huynh

Generalized likelihood ratio (GLR) test statistics are often used in the detection of spatial clustering in case-control and case-population datasets to check for a significantly large proportion of cases within some scanning window. The…

Statistics Theory · Mathematics 2009-11-20 Hock Peng Chan

Large language models (LLMs) show inherent brittleness in their safety mechanisms, as evidenced by their susceptibility to jailbreaking and even non-malicious fine-tuning. This study explores this brittleness of safety alignment by…

Machine Learning · Computer Science 2024-10-28 Boyi Wei , Kaixuan Huang , Yangsibo Huang , Tinghao Xie , Xiangyu Qi , Mengzhou Xia , Prateek Mittal , Mengdi Wang , Peter Henderson

High-dimensional statistical inference with general estimating equations are challenging and remain less explored. In this paper, we study two problems in the area: confidence set estimation for multiple components of the model parameters,…

Methodology · Statistics 2021-04-28 Jinyuan Chang , Song Xi Chen , Cheng Yong Tang , Tong Tong Wu

Mediation analysis aims to assess if, and how, a certain exposure influences an outcome of interest through intermediate variables. This problem has recently gained a surge of attention due to the tremendous need for such analyses in…

Methodology · Statistics 2023-10-19 He Yinqiu , Song Peter X. -K. , Xu Gongjun

We introduce estimation and test procedures through divergence optimization for discrete or continuous parametric models. This approach is based on a new dual representation for divergences. We treat point estimation and tests for simple…

Statistics Theory · Mathematics 2008-12-02 Michel Broniatowski , Amor Keziou

In an instrumental variable model, the score statistic can be bounded for any alternative in parts of the parameter space. These regions involve a constraint on the first-stage regression coefficients and the reduced-form covariance matrix.…

Statistics Theory · Mathematics 2021-09-13 Marcelo J. Moreira , Geert Ridder

Recently, overconfidence in large language models (LLMs) has garnered considerable attention due to its fundamental importance in quantifying the trustworthiness of LLM generation. However, existing approaches prompt the \textit{black box…

Computation and Language · Computer Science 2025-04-29 Adil Bahaj , Hamed Rahimi , Mohamed Chetouani , Mounir Ghogho

The classical likelihood ratio test (LRT) based on the asymptotic chi-squared distribution of the log likelihood is one of the fundamental tools of statistical inference. A recent universal LRT approach based on sample splitting provides…

Methodology · Statistics 2022-11-22 Robin Dunn , Aaditya Ramdas , Sivaraman Balakrishnan , Larry Wasserman

As LLMs are deployed in high-stakes settings, users must judge the correctness of individual responses, often relying on model-generated justifications such as reasoning chains or explanations. Yet, no standard measure exists for whether…

Intuitively, unfamiliarity should lead to lack of confidence. In reality, current algorithms often make highly confident yet wrong predictions when faced with relevant but unfamiliar examples. A classifier we trained to recognize gender is…

Computer Vision and Pattern Recognition · Computer Science 2020-09-09 Zhizhong Li , Derek Hoiem

The likelihood ratio test is widely used in exploratory factor analysis to assess the model fit and determine the number of latent factors. Despite its popularity and clear statistical rationale, researchers have found that when the…

Statistics Theory · Mathematics 2025-01-08 Yinqiu He , Zi Wang , Gongjun Xu

We analyze theoretical properties of the hybrid test for superior predictability. We demonstrate with a simple example that the test may not be pointwise asymptotically of level $\alpha$ at commonly used significance levels and may lead to…

Econometrics · Economics 2021-09-13 Deborah Kim

In this article, we study the hypothesis testing of the blip / net effects of treatments in a treatment sequence. We illustrate that the likelihood ratio test and the score test may suffer from the curse of dimensionality, the null paradox…

Methodology · Statistics 2020-02-04 Xiaoqin Wang , Li Yin

Profile likelihood confidence intervals are a robust alternative to Wald's method if the asymptotic properties of the maximum likelihood estimator are not met. However, the constrained optimization problem defining profile likelihood…

Computation · Statistics 2021-05-10 Samuel M. Fischer , Mark A. Lewis

Data-driven most powerful tests are statistical hypothesis decision-making tools that deliver the greatest power against a fixed null hypothesis among all corresponding data-based tests of a given size. When the underlying data…

Statistics Theory · Mathematics 2023-03-15 Albert Vexler , Alan D. Hutson

In this paper, we propose considering an exact likelihood score (ELS) test for non-inferiority comparison and we derive its test-based confidence interval for the difference between two independent binomial proportions. The p-value for this…

Methodology · Statistics 2022-07-12 Hassan Lakkis , Andrew Lakkis
‹ Prev 1 4 5 6 7 8 10 Next ›