English
Related papers

Related papers: Discriminant analysis in small and large dimension…

200 papers

Due to the increasing recording capability, functional data analysis has become an important research topic. For functional data the study of outlier detection and/or the development of robust statistical procedures has started recently.…

Statistics Theory · Mathematics 2018-04-13 Graciela Boente , Daniela Rodriguez , Mariela Sued

In this paper, the asymptotic distributions of estimators for the regularized functional canonical correlation and variates of the population are derived. The method is based on the possibility of expressing these regularized quantities as…

Statistics Theory · Mathematics 2007-11-29 J. Cupidon , D. S. Gilliam , R. Eubank , F. Ruymgaart

Deep latent-variable models learn representations of high-dimensional data in an unsupervised manner. A number of recent efforts have focused on learning representations that disentangle statistically independent axes of variation by…

In many statistical signal processing applications, the estimation of nuisance parameters and parameters of interest is strongly linked to the resulting performance. Generally, these applications deal with complex data. This paper focuses…

Applications · Statistics 2016-08-24 Melanie Mahot , Philippe Forster , Frederic Pascal , Jean-Philippe Ovarlez

In this paper we study the asymptotic normality in high-dimensional linear regression. We focus on the case where the covariance matrix of the regression variables has a KMS structure, in asymptotic settings where the number of predictors,…

Statistics Theory · Mathematics 2022-05-17 Saulius Jokubaitis , Remigijus Leipus

Discrete random probability measures are a key ingredient of Bayesian nonparametric inferential procedures. A sample generates ties with positive probability and a fundamental object of both theoretical and applied interest is the…

Statistics Theory · Mathematics 2021-01-20 Pierpaolo De Blasi , Ramsés H. Mena , Igor Prünster

We consider estimation of mean and covariance functions of functional snippets, which are short segments of functions possibly observed irregularly on an individual specific subinterval that is much shorter than the entire study interval.…

Methodology · Statistics 2020-06-08 Zhenhua Lin , Jane-Ling Wang

Utilizing covariate information has been a powerful approach to improve the efficiency and accuracy for causal inference, which support massive amount of randomized experiments run on data-driven enterprises. However, state-of-art…

Methodology · Statistics 2023-11-06 Yuhang Wu , Jinghai He , Zeyu Zheng

This paper gives a theoretical analysis of high dimensional linear discrimination of Gaussian data. We study the excess risk of linear discriminant rules. We emphasis on the poor performances of standard procedures in the case when…

Statistics Theory · Mathematics 2010-02-19 Robin Girard

This article conducts a large dimensional study of a simple yet quite versatile classification model, encompassing at once multi-task and semi-supervised learning, and taking into account uncertain labeling. Using tools from random matrix…

Machine Learning · Statistics 2024-02-22 Victor Leger , Romain Couillet

Margin-based classifiers have been popular in both machine learning and statistics for classification problems. Since a large number of classifiers are available, one natural question is which type of classifiers should be used given a…

Machine Learning · Statistics 2021-10-19 Hanwen Huang , Qinglong Yang

In many social, economical, biological and medical studies, one objective is to classify a subject into one of several classes based on a set of variables observed from the subject. Because the probability distribution of the variables is…

Statistics Theory · Mathematics 2011-05-19 Jun Shao , Yazhen Wang , Xinwei Deng , Sijian Wang

Two semimetrics on probability distributions are proposed, given as the sum of differences of expectations of analytic functions evaluated at spatial or frequency locations (i.e, features). The features are chosen so as to maximize the…

Machine Learning · Statistics 2016-10-31 Wittawat Jitkrittum , Zoltan Szabo , Kacper Chwialkowski , Arthur Gretton

We study two-sample variable selection: identifying variables that discriminate between the distributions of two sets of data vectors. Such variables help scientists understand the mechanisms behind dataset discrepancies. Although…

Machine Learning · Statistics 2025-11-06 Kensuke Mitsuzawa , Motonobu Kanagawa , Stefano Bortoli , Margherita Grossi , Paolo Papotti

Linear diffusions are used to model a large number of stochastic processes in physics, including small mechanical and electrical systems perturbed by thermal noise, as well as Brownian particles controlled by electrical and optical forces.…

Statistical Mechanics · Physics 2023-05-10 Johan du Buisson , Hugo Touchette

The assumption of normality in data has been considered in the field of statistical analysis for a long time. However, in many practical situations, this assumption is clearly unrealistic. It has recently been suggested that the use of…

Computation · Statistics 2016-11-25 Reinaldo B. Arellano-Valle , Javier E. Contreras-Reyes

The problem of estimating a linear functional based on observational data is canonical in both the causal inference and bandit literatures. We analyze a broad class of two-stage procedures that first estimate the treatment effect function,…

Statistics Theory · Mathematics 2022-09-28 Wenlong Mou , Martin J. Wainwright , Peter L. Bartlett

High dimensionality comparable to sample size is common in many statistical problems. We examine covariance matrix estimation in the asymptotic framework that the dimensionality $p$ tends to $\infty$ as the sample size $n$ increases.…

Statistics Theory · Mathematics 2007-06-13 Jianqing Fan , Yingying Fan , Jinchi Lv

As the size of modern data sets exceeds the disk and memory capacities of a single computer, machine learning practitioners have resorted to parallel and distributed computing. Given that optimization is one of the pillars of machine…

Machine Learning · Statistics 2019-12-10 Biyi Fang , Diego Klabjan

In this paper, we consider testing the correlation coefficient matrix between two subsets of high-dimensional variables. We produce a test statistic by using the extended cross-data-matrix (ECDM) methodology and show the unbiasedness of…

Methodology · Statistics 2015-03-24 Kazuyoshi Yata , Makoto Aoshima