中文
相关论文

相关论文: Heywood cases in unidimensional factor models and …

200 篇论文

Hierarchical categorical variables often exhibit many levels (high granularity) and many classes within each level (high dimensionality). This may cause overfitting and estimation issues when including such covariates in a predictive model.…

统计方法学 · 统计学 2024-08-20 Paul Wilsens , Katrien Antonio , Gerda Claeskens

Data fusion techniques integrate information from heterogeneous data sources to improve learning, generalization, and decision making across data sciences. In causal inference, these methods leverage rich observational data to improve…

统计方法学 · 统计学 2025-06-02 Quinn Lanners , Cynthia Rudin , Alexander Volfovsky , Harsh Parikh

Different dependence scenarios can arise in multivariate extremes, entailing careful selection of an appropriate class of models. In bivariate extremes, the variables are either asymptotically dependent or are asymptotically independent.…

统计方法学 · 统计学 2015-10-30 Jennifer Wadsworth , Jonathan Tawn , Anthony Davison , Daniel Elton

In causal inference, interference occurs when the treatment of one unit may affect the outcomes of other units. The goal of this work is to serve as a guide to the use of linear outcome modeling for estimating causal effects in settings…

统计方法学 · 统计学 2026-04-01 Eric Tong , Salvador V. Balkus

We encounter variables with little variation often in educational data mining (EDM) due to the demographics of higher education and the questions we ask. Yet, little work has examined how to analyze such data. Therefore, we conducted a…

统计方法学 · 统计学 2022-01-12 Nicholas T. Young , Marcos D. Caballero

In practice functional data are sampled on a discrete set of observation points and often susceptible to noise. We consider in this paper the setting where such data are used as explanatory variables in a regression problem. If the primary…

统计方法学 · 统计学 2021-12-14 Siegfried Hörmann , Fatima Jammoul

The technique of data augmentation (DA) is often used in machine learning for regularization purposes to better generalize under i.i.d. settings. In this work, we present a unifying framework with topics in causal inference to make a case…

机器学习 · 计算机科学 2026-02-02 Uzair Akbar , Niki Kilbertus , Hao Shen , Krikamol Muandet , Bo Dai

Given a large number of covariates $Z$, we consider the estimation of a high-dimensional parameter $\theta$ in an individualized linear threshold $\theta^T Z$ for a continuous variable $X$, which minimizes the disagreement between…

统计理论 · 数学 2019-05-28 Huijie Feng , Yang Ning , Jiwei Zhao

When inferring the causal effect of one variable on another from correlational data, a common practice by professional researchers as well as lay decision makers is to control for some set of exogenous confounding variables. Choosing an…

理论经济学 · 经济学 2023-05-31 Ran Spiegler

The paper is devoted to the consequences of blind random selection of items from different item populations that might be based on completely uncorrelated factors for item inter-correlations and corresponding factor loadings. Based on the…

应用统计 · 统计学 2021-04-26 André Beauducel , Norbert Hilger

Widely used methods for analyzing missing data can be biased in small samples. To understand these biases, we evaluate in detail the situation where a small univariate normal sample, with values missing at random, is analyzed using either…

统计理论 · 数学 2017-03-27 Paul T. von Hippel

Inverse optimization (IO) is used to estimate unknown parameters of an optimization model from observed decisions. In the data-driven context, the estimated parameters are inherently uncertain, yet quantifying this uncertainty has received…

最优化与控制 · 数学 2026-05-26 Timothy C. Y. Chan , Nathan Sandholtz , Nasrin Yousefi

Imbalanced data sets containing much more background than signal instances are very common in particle physics, and will also be characteristic for the upcoming analyses of LHC data. Following up the work presented at ACAT 2008, we use the…

数据分析、统计与概率 · 物理学 2011-08-11 Markward Britsch , Nikolai Gagunashvili , Michael Schmelling

We introduce a new method of estimation of parameters in semiparametric and nonparametric models. The method is based on estimating equations that are $U$-statistics in the observations. The $U$-statistics are based on higher order…

统计方法学 · 统计学 2023-07-14 James Robins , Lingling Li , Rajarshi Mukherjee , Eric Tchetgen Tchetgen , Aad van der Vaart

There has been a lot of work fitting Ising models to multivariate binary data in order to understand the conditional dependency relationships between the variables. However, additional covariates are frequently recorded together with the…

机器学习 · 统计学 2012-09-28 Jie Cheng , Elizaveta Levina , Pei Wang , Ji Zhu

We develop and analyze algorithms for instrumental variable regression by viewing the problem as a conditional stochastic optimization problem. In the context of least-squares instrumental variable regression, our algorithms neither require…

机器学习 · 统计学 2024-05-31 Xuxing Chen , Abhishek Roy , Yifan Hu , Krishnakumar Balasubramanian

In this paper, the asymptotic distributions of estimators for the regularized functional canonical correlation and variates of the population are derived. The method is based on the possibility of expressing these regularized quantities as…

统计理论 · 数学 2007-11-29 J. Cupidon , D. S. Gilliam , R. Eubank , F. Ruymgaart

Linear fixed point equations in Hilbert spaces arise in a variety of settings, including reinforcement learning, and computational methods for solving differential and integral equations. We study methods that use a collection of random…

机器学习 · 计算机科学 2020-12-11 Wenlong Mou , Ashwin Pananjady , Martin J. Wainwright

The Delta method is a classical procedure for quantifying epistemic uncertainty in statistical models, but its direct application to deep neural networks is prevented by the large number of parameters $P$. We propose a low cost variant of…

机器学习 · 计算机科学 2021-03-02 Geir K. Nilsen , Antonella Z. Munthe-Kaas , Hans J. Skaug , Morten Brun

Factor analysis is a critical component of high dimensional biological data analysis. However, modern biological data contain two key features that irrevocably corrupt existing methods. First, these data, which include longitudinal,…

统计方法学 · 统计学 2020-09-24 Chris McKennan