中文
相关论文

相关论文: A regularized MANOVA test for semicontinuous high-…

200 篇论文

When facing multivariate covariates, general semiparametric regression techniques come at hand to propose flexible models that are unexposed to the curse of dimensionality. In this work a semiparametric copula-based estimator for…

统计方法学 · 统计学 2016-03-25 Mickael De Backer , Anouar El Ghouch , Ingrid Van Keilegom

The ratio between two probability density functions is an important component of various tasks, including selection bias correction, novelty detection and classification. Recently, several estimators of this ratio have been proposed. Most…

统计方法学 · 统计学 2014-04-30 Rafael Izbicki , Ann B. Lee , Chad M. Schafer

Data dimensionality informs us about data complexity and sets limit on the structure of successful signal processing pipelines. In this work we revisit and improve the manifold-adaptive Farahmand-Szepesv\'ari-Audibert (FSA) dimension…

In this paper we tackle the ANOVA problem for directional data (with particular emphasis on geological data) by having recourse to the Le Cam methodology usually reserved for linear multivariate analysis. We construct locally and…

统计理论 · 数学 2012-12-07 Christophe Ley , Yvik Swan , Thomas Verdebout

Most normality tests in the literature are performed for scalar and independent samples. Thus, they become unreliable when applied to colored processes, hampering their use in realistic scenarios.We focus on Mardia's multivariate kurtosis,…

统计方法学 · 统计学 2022-03-02 Sara Elbouch , Olivier Michel , Pierre Comon

Two-sample hypothesis testing-determining whether two sets of data are drawn from the same distribution-is a fundamental problem in statistics and machine learning with broad scientific applications. In the context of nonparametric testing,…

机器学习 · 统计学 2026-04-21 Antoine Chatalic , Marco Letizia , Nicolas Schreuder , Lorenzo Rosasco

We propose a new testing procedure of heteroskedasticity in high-dimensional linear regression, where the number of covariates can be larger than the sample size. Our testing procedure is based on residuals of the Lasso. We demonstrate that…

统计理论 · 数学 2022-11-01 Akira Shinkyu

We compare two theoretically distinct approaches to generating artificial (or ``surrogate'') data for testing hypotheses about a given data set. The first and more straightforward approach is to fit a single ``best'' model to the original…

comp-gas · 物理学 2015-06-24 James Theiler , Dean Prichard

While the problem of testing multivariate normality has received considerable attention in the classical low-dimensional setting where the sample size $n$ is much larger than the feature dimension $d$ of the data, there is presently a…

统计方法学 · 统计学 2025-12-23 Xin Bing , Derek Latremouille

In many complex statistical models maximum likelihood estimators cannot be calculated. In the paper we solve this problem using Markov chain Monte Carlo approximation of the true likelihood. In the main result we prove asymptotic normality…

统计理论 · 数学 2018-08-09 Błażej Miasojedow , Wojciech Niemiro , Wojciech Rejchel

We consider the problem of two-sample testing in a semi-supervised setting with abundant unlabeled covariate data. Standard two-sample tests neglect covariate information, which has the potential to significantly boost performance. However,…

机器学习 · 统计学 2026-05-05 Gyumin Lee , Shubhanshu Shekhar , Ilmun Kim

The paper introduces a novel approach to global sensitivity analysis, grounded in the variance-covariance structure of random variables derived from random measures. The proposed methodology facilitates the application of…

统计方法学 · 统计学 2025-10-20 Caleb Deen Bastian , Herschel Rabitz , Grzegorz A Rempala

Modern data analysis depends increasingly on estimating models via flexible high-dimensional or nonparametric machine learning methods, where the identification of structural parameters is often challenging and untestable. In linear…

统计理论 · 数学 2026-01-21 Andrii Babii , Jean-Pierre Florens

Existing feature filters rely on statistical pair-wise dependence metrics to model feature-target relationships, but this approach may fail when the target depends on higher-order feature interactions rather than individual contributions.…

机器学习 · 计算机科学 2025-10-07 Taurai Muvunza , Egor Kraev , Pere Planell-Morell , Alexander Y. Shestopaloff

We propose a likelihood ratio test framework for testing normal mean vectors in high-dimensional data under two common scenarios: the one-sample test and the two-sample test with equal covariance matrices. We derive the test statistics…

统计方法学 · 统计学 2018-09-25 Zongliang Hu , Tiejun Tong , Marc G. Genton

The EM algorithm is a generic tool that offers maximum likelihood solutions when datasets are incomplete with data values missing at random or completely at random. At least for its simplest form, the algorithm can be rewritten in terms of…

统计方法学 · 统计学 2025-09-25 Daniel A. Griffith

Temporal dependence and the resulting autocovariances in time series data can introduce bias into ANOVA test statistics, thereby affecting their size and power. This manuscript accounts for temporal dependence in ANOVA and develops a test…

统计理论 · 数学 2025-09-12 Yunyi Zhang

The regularization approach for variable selection was well developed for a completely observed data set in the past two decades. In the presence of missing values, this approach needs to be tailored to different missing data mechanisms. In…

统计方法学 · 统计学 2017-07-31 Jiwei Zhao , Yang Yang , Yang Ning

We investigate one/two-sample mean tests for high-dimensional compositional data when the number of variables is comparable with the sample size, as commonly encountered in microbiome research. Existing methods mainly focus on max-type test…

统计理论 · 数学 2024-04-15 Qianqian Jiang , Wenbo Li , Zeng Li

We carry out ANOVA comparisons of multiple treatments for longitudinal studies with missing values. The treatment effects are modeled semiparametrically via a partially linear regression which is flexible in quantifying the time effects of…

统计理论 · 数学 2012-11-14 Song Xi Chen , Ping-Shou Zhong