中文
相关论文

相关论文: Hypothesis Testing of One-Sample Mean Vector in Di…

200 篇论文

We consider a number of fundamental statistical and graph problems in the message-passing model, where we have $k$ machines (sites), each holding a piece of data, and the machines want to jointly solve a problem defined on the union of the…

数据结构与算法 · 计算机科学 2013-07-29 David P. Woodruff , Qin Zhang

Distributed high dimensional mean estimation is a common aggregation routine used often in distributed optimization methods. Most of these applications call for a communication-constrained setting where vectors, whose mean is to be…

机器学习 · 统计学 2026-01-28 Harsh Vardhan , Arya Mazumdar

In this paper, we investigate hypothesis testing for the linear combination of mean vectors across multiple populations through the method of random integration. We have established the asymptotic distributions of the test statistics under…

应用统计 · 统计学 2024-03-13 Jianghao Li , Shizhe Hong , Zhenzhen Niu , Zhidong Bai

The divide and conquer method is a common strategy for handling massive data. In this article, we study the divide and conquer method for cubic-rate estimators under the massive data framework. We develop a general theory for establishing…

统计理论 · 数学 2017-04-06 Chengchun Shi , Wenbin Lu , Rui Song

We consider the problem of sparse normal means estimation in a distributed setting with communication constraints. We assume there are $M$ machines, each holding $d$-dimensional observations of a $K$-sparse vector $\mu$ corrupted by…

机器学习 · 统计学 2022-02-15 Chen Amiraz , Robert Krauthgamer , Boaz Nadler

In modern scientific research, massive datasets with huge numbers of observations are frequently encountered. To facilitate the computational process, a divide-and-conquer scheme is often used for the analysis of big data. In such a…

机器学习 · 统计学 2015-05-06 Chen Xu , Yongquan Zhang , Runze Li

The problem of distributed testing against independence with variable-length coding is considered when the \emph{average} and not the \emph{maximum} communication load is constrained as in previous works. The paper characterizes the optimum…

信息论 · 计算机科学 2020-05-19 Sadaf Salehkalaibar , Michele Wigger

Even though a train/test split of the dataset randomly performed is a common practice, could not always be the best approach for estimating performance generalization under some scenarios. The fact is that the usual machine learning…

机器学习 · 计算机科学 2022-09-09 Carlos Catania , Jorge Guerra , Juan Manuel Romero , Gabriel Caffaratti , Martin Marchetta

In this article, we focus on the problem of testing the equality of several high dimensional mean vectors with unequal covariance matrices. This is one of the most important problem in multivariate statistical analysis and there have been…

统计理论 · 数学 2015-04-28 Jiang Hu , Zhidong Bai , Chen Wang , Wei Wang

A popular approach for testing if two univariate random variables are statistically independent consists of partitioning the sample space into bins, and evaluating a test statistic on the binned data. The partition size matters, and the…

统计方法学 · 统计学 2016-04-28 Ruth Heller , Yair Heller , Shachar Kaufman , Barak Brill , Malka Gorfine

This paper studies model checking for general parametric regression models having no dimension reduction structures on the predictor vector. Using any U-statistic type test as an initial test, this paper combines the sample-splitting and…

统计方法学 · 统计学 2023-08-21 Feng Liang , Chuhan Wang , jiaqi Huang , Lixing Zhu

Distributed averaging, or distributed average consensus, is a common method for computing the sample mean of the data dispersed among the nodes of a network in a decentralized manner. By iteratively exchanging messages with neighbors, the…

信息论 · 计算机科学 2017-10-26 Ryan Pilgrim , Junan Zhu , Dror Baron , Waheed U. Bajwa

We consider the problem where $n$ clients transmit $d$-dimensional real-valued vectors using $d(1+o(1))$ bits each, in a manner that allows the receiver to approximately reconstruct their mean. Such compression problems naturally arise in…

机器学习 · 计算机科学 2021-12-17 Shay Vargaftik , Ran Ben Basat , Amit Portnoy , Gal Mendelson , Yaniv Ben-Itzhak , Michael Mitzenmacher

A common approach to statistical learning with big-data is to randomly split it among $m$ machines and learn the parameter of interest by averaging the $m$ individual estimates. In this paper, focusing on empirical risk minimization, or…

机器学习 · 统计学 2016-06-14 Jonathan Rosenblatt , Boaz Nadler

This paper considers distributed M-estimation under heterogeneous distributions among distributed data blocks. A weighted distributed estimator is proposed to improve the efficiency of the standard "Split-And-Conquer" (SaC) estimator for…

统计理论 · 数学 2022-09-15 Jia Gu , Songxi Chen

Large data sets often require performing distributed statistical estimation, with a full data set split across multiple machines and limited communication between machines. To study such scenarios, we define and study some refinements of…

信息论 · 计算机科学 2014-06-24 John C. Duchi , Michael I. Jordan , Martin J. Wainwright , Yuchen Zhang

This paper considers testing linear hypotheses of a set of mean vectors with unequal covariance matrices in large dimensional setting. The problem of testing the hypothesis $H_0 : \sum_{i=1}^q \beta_i \bmu_i =\bmu_0 $ for a given vector…

统计方法学 · 统计学 2015-12-22 Dandan Jiang

We study the fundamental problem of Principal Component Analysis in a statistical distributed setting in which each machine out of $m$ stores a sample of $n$ points sampled i.i.d. from a single unknown distribution. We study algorithms for…

机器学习 · 计算机科学 2017-02-28 Dan Garber , Ohad Shamir , Nathan Srebro

We investigate the problem of jointly testing two hypotheses and estimating a random parameter based on data that is observed sequentially by sensors in a distributed network. In particular, we assume the data to be drawn from a Gaussian…

信号处理 · 电气工程与系统科学 2020-03-04 Dominik Reinhard , Michael Fauß , Abdelhak M. Zoubir

It is not unusual for a data analyst to encounter data sets distributed across several computers. This can happen for reasons such as privacy concerns, efficiency of likelihood evaluations, or just the sheer size of the whole data set. This…

统计计算 · 统计学 2018-05-22 Randy C. S. Lai , J. Hannig , Thomas C. M. Lee