中文
相关论文

相关论文: Scalable High-Dimensional Multivariate Linear Regr…

200 篇论文

Distributed statistical learning problems arise commonly when dealing with large datasets. In this setup, datasets are partitioned over machines, which compute locally, and communicate short messages. Communication is often the bottleneck.…

统计理论 · 数学 2022-10-25 Edgar Dobriban , Yue Sheng

We consider the problem of computing a Gaussian approximation to the posterior distribution of a parameter given a large number N of observations and a Gaussian prior, when the dimension of the parameter d is also large. To address this…

数据结构与算法 · 计算机科学 2023-03-28 Marc Lambert , Silvère Bonnabel , Francis Bach

The demand of computational resources for the modeling process increases as the scale of the datasets does, since traditional approaches for regression involve inverting huge data matrices. The main problem relies on the large data size,…

统计方法学 · 统计学 2023-07-06 Vasilis Chasiotis , Dimitris Karlis

Previous work has demonstrated the feasibility and value of conducting distributed regression analysis (DRA), a privacy-protecting analytic method that performs multivariable-adjusted regression analysis with only summary-level information…

统计计算 · 统计学 2018-08-08 Qoua L. Her , Yury Vilk , Jessica Young , Zilu Zhang , Jessica M. Malenfant , Sarah Malek , Sengwee Toh

In many social, economical, biological and medical studies, one objective is to classify a subject into one of several classes based on a set of variables observed from the subject. Because the probability distribution of the variables is…

统计理论 · 数学 2011-05-19 Jun Shao , Yazhen Wang , Xinwei Deng , Sijian Wang

Stochastic optimization algorithms update models with cheap per-iteration costs sequentially, which makes them amenable for large-scale data analysis. Such algorithms have been widely studied for structured sparse models where the sparsity…

机器学习 · 计算机科学 2019-05-10 Baojian Zhou , Feng Chen , Yiming Ying

Rate-splitting multiple access (RSMA) has been proven as an effective communication scheme for 5G and beyond. However, current approaches to RSMA resource management require complicated iterative algorithms, which cannot meet the stringent…

信息论 · 计算机科学 2024-11-07 Hanwen Zhang , Mingzhe Chen , Alireza Vahid , Feng Ye , Haijian Sun

Hierarchical models are a powerful tool for high-throughput data with a small to moderate number of replicates, as they allow sharing information across units of information, for example, genes. We propose two such models and show its…

应用统计 · 统计学 2009-10-09 David Rossell

A Two-Stage approach enables researchers to make optimal non-linear predictions via Generalized Ridge Regression using models that contain two or more x-predictor variables and make only realistic minimal assumptions. The optimal regression…

统计方法学 · 统计学 2023-07-11 Robert L. Obenchain

We consider the problem of breaking a multivariate (vector) time series into segments over which the data is well explained as independent samples from a Gaussian distribution. We formulate this as a covariance-regularized maximum…

最优化与控制 · 数学 2018-04-30 David Hallac , Peter Nystrup , Stephen Boyd

We tackle the challenges of modeling high-dimensional data sets, particularly those with latent low-dimensional structures hidden within complex, non-linear, and noisy relationships. Our approach enables a seamless integration of concepts…

机器学习 · 统计学 2025-03-17 Zichuan Guo , Mihai Cucuringu , Alexander Y. Shestopaloff

Conventional feature selection algorithms applied to Pseudo Time-Series (PTS) data, which consists of observations arranged in sequential order without adhering to a conventional temporal dimension, often exhibit impractical computational…

机器学习 · 计算机科学 2024-03-14 Mohammad Rahman , Manzur Murshed , Shyh Wei Teng , Manoranjan Paul

This paper studies the distributed adaptiveestimation problems for stochastic large regression modelswith an infinite number of parameters. By constructing a re-cursive local cost function, we propose a novel distributedrecursive least…

系统与控制 · 电气工程与系统科学 2026-04-29 Die Gan , Siyu Xie , Zhixin Liu , Xuebo Zhang

Sparsity-constrained optimization has wide applicability in machine learning, statistics, and signal processing problems such as feature selection and compressive Sensing. A vast body of work has studied the sparsity-constrained…

机器学习 · 统计学 2013-07-17 Sohail Bahmani , Bhiksha Raj , Petros Boufounos

Factor Analysis based on multivariate $t$ distribution ($t$fa) is a useful robust tool for extracting common factors on heavy-tailed or contaminated data. However, $t$fa is only applicable to vector data. When $t$fa is applied to matrix…

机器学习 · 统计学 2024-01-05 Xuan Ma , Jianhua Zhao , Changchun Shang , Fen Jiang , Philip L. H. Yu

Scaling multinomial logistic regression to datasets with very large number of data points and classes is challenging. This is primarily because one needs to compute the log-partition function on every data point. This makes distributing the…

Maximum weight matching is one of the most fundamental combinatorial optimization problems with a wide range of applications in data mining and bioinformatics. Developing distributed weighted matching algorithms is challenging due to the…

分布式、并行与集群计算 · 计算机科学 2019-06-06 Sepehr Assadi , MohammadHossein Bateni , Vahab Mirrokni

Modern datasets arising from social media, genomics, and biomedical informatics are often heterogeneous and (ultra) high-dimensional, creating substantial challenges for conventional modeling techniques. Quantile regression (QR) not only…

统计方法学 · 统计学 2026-01-07 Hanqing Wu , Jonas Wallin , Iuliana Ionita-Laza

We analyse the learning performance of Distributed Gradient Descent in the context of multi-agent decentralised non-parametric regression with the square loss function when i.i.d. samples are assigned to agents. We show that if agents hold…

机器学习 · 统计学 2019-11-14 Dominic Richards , Patrick Rebeschini

Several learning applications require solving high-dimensional regression problems where the relevant features belong to a small number of (overlapping) groups. For very large datasets and under standard sparsity constraints, hard…

机器学习 · 统计学 2016-05-30 Prateek Jain , Nikhil Rao , Inderjit Dhillon