English
Related papers

Related papers: Covariate dimension reduction for survival data vi…

200 papers

The high-dimensional data setting, in which p >> n, is a challenging statistical paradigm that appears in many real-world problems. In this setting, learning a compact, low-dimensional representation of the data can substantially help…

Machine Learning · Computer Science 2018-08-07 Micol Marchetti-Bowick , Benjamin J. Lengerich , Ankur P. Parikh , Eric P. Xing

Gaussian graphical models are widely utilized to infer and visualize networks of dependencies between continuous variables. However, inferring the graph is difficult when the sample size is small compared to the number of variables. To…

Statistics Theory · Mathematics 2016-09-30 Emilie Devijver , Mélina Gallopin

While covariance matrices have been widely studied in many scientific fields, relatively limited progress has been made on estimating conditional covariances that permits a large covariance matrix to vary with high-dimensional subject-level…

Methodology · Statistics 2025-05-28 Rakheon Kim , Jingfei Zhang

We derive improved regression and classification rates for support vector machines using Gaussian kernels under the assumption that the data has some low-dimensional intrinsic structure that is described by the box-counting dimension. Under…

Statistics Theory · Mathematics 2021-04-08 Thomas Hamm , Ingo Steinwart

In medical and biological research, longitudinal data and survival data types are commonly seen. Traditional statistical models mostly consider to deal with either of the data types, such as linear mixed models for longitudinal data, and…

Methodology · Statistics 2021-07-12 Jizi Shangguan

The Cox proportional hazards model is the most widely used regression model in univariate survival analysis. Extensions of the Cox model to bivariate survival data, however, remain scarce. We propose two novel extensions based on a…

Methodology · Statistics 2025-11-12 Yael Travis-Lumer , Micha Mandel , Ido Didi Fabian , Rebecca A. Betensky , Malka Gorfine

Graphical models are commonly used to represent conditional dependence relationships between variables. There are multiple methods available for exploring them from high-dimensional data, but almost all of them rely on the assumption that…

Machine Learning · Statistics 2020-04-22 Tianxi Li , Cheng Qian , Elizaveta Levina , Ji Zhu

Gaussian Graphical Models (GGMs) are popular tools for studying network structures. However, many modern applications such as gene network discovery and social interactions analysis often involve high-dimensional noisy data with outliers or…

Machine Learning · Statistics 2015-10-30 Eunho Yang , Aurélie C. Lozano

Gaussian graphical models (GGM) have been widely used in many high-dimensional applications ranging from biological and financial data to recommender systems. Sparsity in GGM plays a central role both statistically and computationally.…

Machine Learning · Statistics 2014-06-12 Zhaoshi Meng , Brian Eriksson , Alfred O. Hero

We consider estimation of a deterministic unknown parameter vector in a linear model with non-Gaussian noise. In the Gaussian case, dimensionality reduction via a linear matched filter provides a simple low dimensional sufficient statistic…

Applications · Statistics 2013-11-05 Jakob Vovnoboy , Ami Wiesel

In applications involving ordinal predictors, common approaches to reduce dimensionality are either extensions of unsupervised techniques such as principal component analysis, or variable selection procedures that rely on modeling the…

Statistics Theory · Mathematics 2017-10-13 Liliana Forzani , Rodrigo García Arancibia , Pamela Llop , Diego Tomassi

This paper gives a theoretical analysis of high dimensional linear discrimination of Gaussian data. We study the excess risk of linear discriminant rules. We emphasis on the poor performances of standard procedures in the case when…

Statistics Theory · Mathematics 2010-02-19 Robin Girard

The covariance matrix plays a fundamental role in many modern exploratory and inferential statistical procedures, including dimensionality reduction, hypothesis testing, and regression. In low-dimensional regimes, where the number of…

Methodology · Statistics 2024-11-12 Philippe Boileau , Nima S. Hejazi , Mark J. van der Laan , Sandrine Dudoit

High-dimensional linear and nonlinear models have been extensively used to identify associations between response and explanatory variables. The variable selection problem is commonly of interest in the presence of massive and complex data.…

Methodology · Statistics 2017-08-10 Vitara Pungpapong , Min Zhang , Dabao Zhang

Recently emerging large-scale biomedical data pose exciting opportunities for scientific discoveries. However, the ultrahigh dimensionality and non-negligible measurement errors in the data may create difficulties in estimation. There are…

Methodology · Statistics 2022-10-28 Xin Ma , Suprateek Kundu

The paper considers variable selection in linear regression models where the number of covariates is possibly much larger than the number of observations. High dimensionality of the data brings in many complications, such as (possibly…

Methodology · Statistics 2016-11-29 Haeran Cho , Piotr Fryzlewicz

In this paper, we study the model selection and structure specification for the generalised semi-varying coefficient models (GSVCMs), where the number of potential covariates is allowed to be larger than the sample size. We first propose a…

Statistics Theory · Mathematics 2015-10-30 Degui Li , Yuan Ke , Wenyang Zhang

Analysis of high-dimensional data is currently a popular field of research, thanks to many applications e.g. in genetics (DNA data in genomewide association studies), spectrometry or web analysis. At the same time, the type of problems that…

Methodology · Statistics 2018-05-25 Jozef Jakubik

In many important statistical analyses, the number of covariates $p$ often exceeds the data size $n$, a regime commonly referred to as high-dimensional. While considerable progress has been made in high-dimensional regression under the…

Methodology · Statistics 2026-05-29 Herman Tesso , Georges Nguefack-Tsague

The paper is motivated from clustering problem in high-throughput mixed datasets. Clustering of such datasets can provide much insight into biological associations. An open problem in this context is to simultaneously cluster…

Methodology · Statistics 2018-08-15 Chetkar Jha
‹ Prev 1 4 5 6 7 8 10 Next ›