中文
相关论文

相关论文: Improved LM Test for Robust Model Specification Se…

200 篇论文

Large language models (LLMs) are increasingly deployed for tabular question answering, yet calibration on structured data is largely unstudied. This paper presents the first systematic comparison of five confidence estimation methods across…

计算与语言 · 计算机科学 2026-04-15 Lukas Voss

In semi-supervised learning, the prevailing understanding suggests that observing additional unlabeled samples improves estimation accuracy for linear parameters only in the case of model misspecification. In this work, we challenge such a…

统计方法学 · 统计学 2025-09-03 Kai Chen , Yuqian Zhang

This paper studies the high-dimensional mixed linear regression (MLR) where the output variable comes from one of the two linear regression models with an unknown mixing proportion and an unknown covariance structure of the random…

统计方法学 · 统计学 2020-11-10 Linjun Zhang , Rong Ma , T. Tony Cai , Hongzhe Li

We introduce a simple diagnostic test for assessing the overall or partial goodness of fit of a linear causal model with errors being independent of the covariates. In particular, we consider situations where hidden confounding is…

统计方法学 · 统计学 2023-03-06 Christoph Schultheiss , Peter Bühlmann , Ming Yuan

Calibration, the practice of choosing the parameters of a structural model to match certain empirical moments, can be viewed as minimum distance estimation. Existing standard error formulas for such estimators require a consistent estimate…

计量经济学 · 经济学 2024-06-19 Matthew D. Cocci , Mikkel Plagborg-Møller

Structural equation models (SEMs) are fundamental to causal mediation pathway discovery. However, traditional SEM approaches often rely on \emph{ad hoc} model specifications when handling complex data structures such as mixed data types or…

统计方法学 · 统计学 2025-10-22 Canyi Chen , Ritoban Kundu , Wei Hao , Peter X. -K. Song

This paper considers errors-in-variables models in a high-dimensional setting where the number of covariates can be much larger than the sample size, and there are only a small number of non-zero covariates. The presence of measurement…

统计方法学 · 统计学 2018-09-03 Linh Nghiem , Cornelis Potgieter

A growing body of work has been querying LLMs with political questions to evaluate their potential biases. However, this probing method has limited stability, making comparisons between models unreliable. In this paper, we argue that LLMs…

计算与语言 · 计算机科学 2025-06-30 Patrick Haller , Jannis Vamvas , Rico Sennrich , Lena A. Jäger

There is ongoing debate about whether large language models (LLMs) can serve as substitutes for human participants in survey and experimental research. While recent work in fields such as marketing and psychology has explored the potential…

计算与语言 · 计算机科学 2025-12-30 Steven Wang , Kyle Hunt , Shaojie Tang , Kenneth Joseph

We study the problem of testing whether the missing values of a potentially high-dimensional dataset are Missing Completely at Random (MCAR). We relax the problem of testing MCAR to the problem of testing the compatibility of a collection…

统计理论 · 数学 2024-12-13 Alberto Bordino , Thomas B. Berrett

Survey researchers face two key challenges: the rising costs of probability samples and missing data (e.g., non-response or attrition), which can undermine inference and increase the use of convenience samples. Recent work explores using…

计算机与社会 · 计算机科学 2025-09-30 Tobias Holtdirk , Dennis Assenmacher , Arnim Bleier , Claudia Wagner

Large language models (LLMs) struggle with compositional generalisation, limiting their ability to systematically combine learned components to interpret novel inputs. While architectural modifications, fine-tuning, and data augmentation…

计算与语言 · 计算机科学 2025-05-21 Nura Aljaafari , Danilo S. Carvalho , André Freitas

Regularized linear discriminant analysis (RLDA) is a widely used tool for classification and dimensionality reduction, but its performance in high-dimensional scenarios is inconsistent. Existing theoretical analyses of RLDA often lack clear…

机器学习 · 统计学 2025-07-23 Yonghan Zhang , Zhangni Pu , Lu Yan , Jiang Hu

In complex engineering systems, the dependencies among components or development activities are often modeled and analyzed using Design Structure Matrix (DSM). Reorganizing elements within a DSM to minimize feedback loops and enhance…

计算工程、金融与科学 · 计算机科学 2026-04-07 Shuo Jiang , Min Xie , Jianxi Luo

The quality of meeting summaries generated by natural language generation (NLG) systems is hard to measure automatically. Established metrics such as ROUGE and BERTScore have a relatively low correlation with human judgments and fail to…

计算与语言 · 计算机科学 2025-02-19 Frederic Kirstein , Terry Ruas , Bela Gipp

Empirical modelling often aims for the simplest model consistent with the data. A new technique is presented which quantifies the consistency of the model dynamics as a function of location in state space. As is well-known, traditional…

混沌动力学 · 物理学 2009-11-10 Patrick E. McSharry , Leonard A. Smith

Linear mixed models (LMMs) have emerged as the method of choice for confounded genome-wide association studies. However, the performance of LMMs in non-randomly ascertained case-control studies deteriorates with increasing sample size. We…

基因组学 · 定量生物学 2016-02-23 Omer Weissbrod , Christoph Lippert , Dan Geiger , David Heckerman

Relying on recent advances in statistical estimation of covariance distances based on random matrix theory, this article proposes an improved covariance and precision matrix estimation for a wide family of metrics. The method is shown to…

机器学习 · 统计学 2021-02-03 Malik Tiomoko , Florent Bouchard , Guillaume Ginholac , Romain Couillet

This paper presents the current state of a work in progress, whose objective is to better understand the effects of factors that significantly influence the performance of Latent Semantic Analysis (LSA). A difficult task, which consists in…

机器学习 · 计算机科学 2009-12-10 Alain Lifchitz , Sandra Jhean-Larose , Guy Denhière

Metamodels, or the regression analysis of Monte Carlo simulation results, provide a powerful tool to summarize simulation findings. However, an underutilized approach is the multilevel metamodel (MLMM) that accounts for the dependent data…

统计方法学 · 统计学 2025-11-21 Joshua Gilbert , Luke Miratrix