English
Related papers

Related papers: Regression and Dimension Reduction for Multivariat…

200 papers

In the social sciences, small- to medium-scale datasets are common, and linear regression is canonical. In privacy-aware settings, much work has focused on differentially private (DP) linear regression, but mostly on point estimation with…

Machine Learning · Computer Science 2026-03-31 Shurong Lin , Aleksandra Slavković , Deekshith Reddy Bhoomireddy

We consider the scenario where one observes an outcome variable and sets of features from multiple assays, all measured on the same set of samples. One approach that has been proposed for dealing with this type of data is ``sparse multiple…

Quantitative Methods · Quantitative Biology 2014-01-24 Samuel M. Gross , Robert Tibshirani

This research is motivated by discovering and underpinning genetic causes for the progression of a bilateral eye disease, Age-related Macular Degeneration (AMD), of which the primary outcomes, progression times to late-AMD, are bivariate…

Methodology · Statistics 2019-08-21 Tao Sun , Ying Ding

Methods for global measurement of transcript abundance such as microarrays and RNA-Seq generate datasets in which the number of measured features far exceeds the number of observations. Extracting biologically meaningful and experimentally…

Methodology · Statistics 2022-06-22 Lei Ding , Gabriel E. Zentner , Daniel J. McDonald

Use of copula for the purpose of modeling dependence has been receiving considerable attention in recent times. On the other hand, search for multivariate copulas with desirable dependence properties also is an important area of research.…

Methodology · Statistics 2025-02-18 Subhajit Chattopadhyay

We propose a semi-supervised generative model, SeGMA, which learns a joint probability distribution of data and their classes and which is implemented in a typical Wasserstein auto-encoder framework. We choose a mixture of Gaussians as a…

Machine Learning · Computer Science 2020-08-28 Marek Śmieja , Maciej Wołczyk , Jacek Tabor , Bernhard C. Geiger

Principal component regression (PCR) is a widely used two-stage procedure: principal component analysis (PCA), followed by regression in which the selected principal components are regarded as new explanatory variables in the model. Note…

Machine Learning · Statistics 2018-04-03 Shuichi Kawano , Hironori Fujisawa , Toyoyuki Takada , Toshihiko Shiroishi

"For how many days during the past 30 days was your mental health not good?" The responses to this question measure self-reported mental health and can be linked to important covariates in the National Health and Nutrition Examination…

Methodology · Statistics 2021-10-14 Daniel R. Kowal , Bohan Wu

Clinical patient records are an example of high-dimensional data that is typically collected from disparate sources and comprises of multiple likelihoods with noisy as well as missing values. In this work, we propose an unsupervised…

Machine Learning · Statistics 2021-04-21 Siddharth Ramchandran , Miika Koskinen , Harri Lähdesmäki

We study a distributed consensus-based stochastic gradient descent (SGD) algorithm and show that the rate of convergence involves the spectral properties of two matrices: the standard spectral gap of a weight matrix from the network…

Optimization and Control · Mathematics 2016-09-02 Avleen S. Bijral , Anand D. Sarwate , Nathan Srebro

Multi-task regression attempts to exploit the task similarity in order to achieve knowledge transfer across related tasks for performance improvement. The application of Gaussian process (GP) in this scenario yields the non-parametric yet…

Machine Learning · Statistics 2021-09-21 Haitao Liu , Jiaqi Ding , Xinyu Xie , Xiaomo Jiang , Yusong Zhao , Xiaofang Wang

Recently, high-dimensional heterogeneous data have attracted a lot of attention and discussion. Under heterogeneity, semiparametric regression is a popular choice to model data in statistics. In this paper, we take advantages of expectile…

Statistics Theory · Mathematics 2019-08-20 Jun Zhao , Guan'ao Yan , Yi Zhang

With the growing prevalence of diabetes and the associated public health burden, it is crucial to identify modifiable factors that could improve patients' glycemic control. In this work, we seek to examine associations between medication…

Applications · Statistics 2025-12-22 Alexander Coulter , Rashmi N. Aurora , Naresh M. Punjabi , Irina Gaynanova

We develop a method to decompose causal effects on a social network into an indirect effect mediated by the network, and a direct effect independent of the social network. To handle the complexity of network structures, we assume that…

Methodology · Statistics 2025-03-07 Alex Hayes , Mark M. Fredrickson , Keith Levin

With the evolution of single-cell RNA sequencing techniques into a standard approach in genomics, it has become possible to conduct cohort-level causal inferences based on single-cell-level measurements. However, the individual gene…

Methodology · Statistics 2025-04-23 Jin-Hong Du , Zhenghao Zeng , Edward H. Kennedy , Larry Wasserman , Kathryn Roeder

This work considers estimation and forecasting in a multivariate, possibly high-dimensional count time series model constructed from a transformation of a latent Gaussian dynamic factor series. The estimation of the latent model parameters…

Methodology · Statistics 2025-04-07 Younghoon Kim , Marie-Christine Düker , Zachary F. Fisher , Vladas Pipiras

Missing data is a common issue in various fields such as medicine, social sciences, and natural sciences, and it poses significant challenges for accurate statistical analysis. Although numerous imputation methods have been proposed to…

Methodology · Statistics 2025-07-23 Seongmin Kim , Jeunghun Oh , Hungkuk Ko , Jeongmin Park , Jaeyong Lee

Variable selection and dimension reduction are two commonly adopted approaches for high-dimensional data analysis, but have traditionally been treated separately. Here we propose an integrated approach, called sparse gradient learning…

Machine Learning · Statistics 2010-07-02 Gui-Bo Ye , Xiaohui Xie

With the rapid advances of data acquisition techniques, spatio-temporal data are becoming increasingly abundant in a diverse array of disciplines. Here we develop spatio-temporal regression methodology for analyzing large amounts of…

Methodology · Statistics 2021-12-01 Ting Fung Ma , Fangfang Wang , Jun Zhu , Anthony R. Ives , Katarzyna E. Lewińska

We face network data from various sources, such as protein interactions and online social networks. A critical problem is to model network interactions and identify latent groups of network nodes. This problem is challenging due to many…

Machine Learning · Computer Science 2012-02-20 Feng Yan , Zenglin Xu , Yuan , Qi