English
Related papers

Related papers: Regression and Dimension Reduction for Multivariat…

200 papers

We investigate robust linear regression where data may be contaminated by an oblivious adversary, i.e., an adversary than may know the data distribution but is otherwise oblivious to the realizations of the data samples. This model has been…

Machine Learning · Computer Science 2022-02-07 Tom Norman , Nir Weinberger , Kfir Y. Levy

In this paper, we propose a general subgroup analysis framework based on semiparametric additive mixed effect models in longitudinal analysis, which can identify subgroups on each covariate and estimate the corresponding regression…

Methodology · Statistics 2021-12-02 Xiaolin Bo , Weiping Zhang

In this article we focus on dynamic network data which describe interactions among a fixed population through time. We model this data using the latent space framework, in which the probability of a connection forming is expressed as a…

Methodology · Statistics 2021-12-21 Kathryn Turnbull , Christopher Nemeth , Matthew Nunes , Tyler McCormick

We introduce Adaptive Subspace PCA (AS-PCA), a framework for principal component analysis of random elements in a general separable Hilbert space. AS-PCA projects the covariance operator onto a data-adaptive finite-dimensional subspace…

Statistics Theory · Mathematics 2026-03-24 Xinyi Li , Margaret Hoch , Michael R. Kosorok

Multivariate data that combine binary, categorical, count and continuous outcomes are common in the social and health sciences. We propose a semiparametric Bayesian latent variable model for multivariate data of arbitrary type that does not…

Applications · Statistics 2014-01-14 Jonathan Gruhl , Elena A. Erosheva , Paul K. Crane

We discuss semiparametric regression when only the ranks of responses are observed. The model is $Y_i = F (\mathbf{x}_i'{\boldsymbol\beta}_0 + \varepsilon_i)$, where $Y_i$ is the unobserved response, $F$ is a monotone increasing function,…

Applications · Statistics 2016-02-25 Michael C. Donohue , Anthony C. Gamst , Robert A. Rissman , Ian Abramson

Unobserved confounding is one of the main challenges when estimating causal effects. We propose a causal reduction method that, given a causal model, replaces an arbitrary number of possibly high-dimensional latent confounders with a single…

Machine Learning · Statistics 2023-02-24 Maximilian Ilse , Patrick Forré , Max Welling , Joris M. Mooij

Theoretically understanding stochastic gradient descent (SGD) in overparameterized models has led to the development of several optimization algorithms that are widely used in practice today. Recent work by~\citet{zou2021benign} provides…

Machine Learning · Computer Science 2025-06-19 Alexandru Meterez , Depen Morwani , Costin-Andrei Oncescu , Jingfeng Wu , Cengiz Pehlevan , Sham Kakade

Gaussian processes (GPs) have gained popularity as flexible machine learning models for regression and function approximation with an in-built method for uncertainty quantification. However, GPs suffer when the amount of training data is…

Machine Learning · Statistics 2025-11-26 Jonas Latz , Aretha L. Teckentrup , Simon Urbainczyk

Probabilistic principal component analysis (PPCA) seeks a low dimensional representation of a data set in the presence of independent spherical Gaussian noise, Sigma = (sigma^2)*I. The maximum likelihood solution for the model is an…

Machine Learning · Statistics 2011-06-23 Alfredo A. Kalaitzis , Neil D. Lawrence

Existing partial sequence labeling models mainly focus on max-margin framework which fails to provide an uncertainty estimation of the prediction. Further, the unique ground truth disambiguation strategy employed by these models may include…

Machine Learning · Computer Science 2022-09-21 Xiaolei Lu , Tommy W. S. Chow

Covariance matrix outcomes arise naturally in neuroimaging experiments to study brain functional connectivity. It is also of interest to understand how brain network organization varies with subject-level covariates. Existing covariance…

Methodology · Statistics 2026-05-08 Michelle Murphy Green , Xi Luo , Brian S. Caffo , Yi Zhao

Missing data imputation forms the first critical step of many data analysis pipelines. The challenge is greatest for mixed data sets, including real, Boolean, and ordinal data, where standard techniques for imputation fail basic sanity…

Methodology · Statistics 2020-06-17 Yuxuan Zhao , Madeleine Udell

Gaussian graphical models (GGMs) are widely used for statistical modeling, because of ease of inference and the ubiquitous use of the normal distribution in practical approximations. However, they are also known for their limited modeling…

Machine Learning · Statistics 2016-11-22 Qinliang Su , Xuejun Liao , Chunyuan Li , Zhe Gan , Lawrence Carin

Modern mobile health (mHealth) assessment combines self-reported measures of participants' health experiences with passively collected health behavior data throughout the day. These data are collected across multiple measurement scales,…

Methodology · Statistics 2026-03-13 Debangan Dey , Rahul Ghosal , Kathleen Merikangas , Vadim Zipunnikov

Principal component analysis (PCA) is commonly used in genetics to infer and visualize population structure and admixture between populations. PCA is often interpreted in a way similar to inferred admixture proportions, where it is assumed…

Methodology · Statistics 2023-02-10 Jan van Waaij , Song Li , Genís Garcia-Erill , Anders Albrechtsen , Carsten Wiuf

Scientific and engineering problems often require the use of artificial intelligence to aid understanding and the search for promising designs. While Gaussian processes (GP) stand out as easy-to-use and interpretable learners, they have…

Machine Learning · Computer Science 2021-07-01 Liwei Wang , Suraj Yerramilli , Akshay Iyer , Daniel Apley , Ping Zhu , Wei Chen

The early detection of Alzheimer's disease (AD) requires an understanding of the relationships between a wide range of features. Conditional independencies and partial correlations are suitable measures for these relationships, because they…

Applications · Statistics 2025-01-22 Lucas Vogels , Reza Mohammadi , Marit Schoonhoven , S. Ilker Birbil , Martin Dyrba

Retrieval-Augmented Generation (RAG) systems for biomedical literature are typically evaluated using ranking metrics like Mean Reciprocal Rank (MRR), which measure how well the system identifies the single most relevant chunk. We argue that…

Artificial Intelligence · Computer Science 2026-03-25 Pouria Mortezaagha , Arya Rahgozar

Understanding and developing a correlation measure that can detect general dependencies is not only imperative to statistics and machine learning, but also crucial to general scientific discovery in the big data age. In this paper, we…

Machine Learning · Statistics 2024-06-27 Cencheng Shen , Carey E. Priebe , Joshua T. Vogelstein