English
Related papers

Related papers: Backfitting for large scale crossed random effects…

200 papers

Large-scale sequential data is often exposed to some degree of inhomogeneity in the form of sudden changes in the parameters of the data-generating process. We consider the problem of detecting such structural changes in a high-dimensional…

Methodology · Statistics 2016-01-15 Florencia Leonardi , Peter Bühlmann

Ordinary least-squares (OLS) estimators for a linear model are very sensitive to unusual values in the design space or outliers among y values. Even one single atypical value may have a large effect on the parameter estimates. This article…

Methodology · Statistics 2014-04-28 Chun Yu , Weixin Yao , Xue Bai

Generalized Linear Models are routinely used in data analysis. The classical procedures for estimation are based on Maximum Likelihood and it is well known that the presence of outliers can have a large impact on this estimator. Robust…

Computation · Statistics 2017-10-02 Marina Valdora , Claudio Agostinelli , Victor J. Yohai

Piecewise deterministic Markov processes provide scalable methods for sampling from the posterior distributions in big data settings by admitting principled sub-sampling strategies that do not bias the output. An important example is the…

Computation · Statistics 2025-08-04 Sanket Agrawal , Joris Bierkens , Gareth O. Roberts

This paper investigates estimation and inference for average treatment effects in completely randomized experiments when researchers observe potentially many covariates. Within Neyman's (1923) design-based framework, allowing the number of…

Econometrics · Economics 2025-11-20 Harold D Chiang , Yukitoshi Matsushita , Taisuke Otsu

We provide an algorithm for properly learning mixtures of two single-dimensional Gaussians without any separability assumptions. Given $\tilde{O}(1/\varepsilon^2)$ samples from an unknown mixture, our algorithm outputs a mixture that is…

Data Structures and Algorithms · Computer Science 2014-05-20 Constantinos Daskalakis , Gautam Kamath

The presence of groups containing high leverage outliers makes linear regression a difficult problem due to the masking effect. The available high breakdown estimators based on Least Trimmed Squares often do not succeed in detecting masked…

Computation · Statistics 2011-03-23 L. Pitsoulis , G. Zioutas

Training Gaussian process-based models typically involves an $ O(N^3)$ computational bottleneck due to inverting the covariance matrix. Popular methods for overcoming this matrix inversion problem cannot adequately model all types of latent…

Machine Learning · Statistics 2020-03-04 Michael Minyi Zhang , Sinead A. Williamson

In this paper, we study the problem of mean estimation under 1-bit communication constraints. We propose a novel adaptive mean estimator based solely on randomized threshold queries, where each 1-bit outcome indicates whether a given sample…

Machine Learning · Statistics 2026-05-25 Ivan Lau , Jonathan Scarlett

In this paper, we propose differentially private algorithms for the problem of stochastic linear bandits in the central, local and shuffled models. In the central model, we achieve almost the same regret as the optimal non-private…

Machine Learning · Computer Science 2022-07-08 Osama A. Hanna , Antonious M. Girgis , Christina Fragouli , Suhas Diggavi

Gaussian graphical modeling has been widely used to explore various network structures, such as gene regulatory networks and social networks. We often use a penalized maximum likelihood approach with the $L_1$ penalty for learning a…

Methodology · Statistics 2017-06-13 Kei Hirose , Hironori Fujisawa , Jun Sese

Bernstein's condition is a key assumption that guarantees fast rates in machine learning. For example, the Gibbs algorithm with prior $\pi$ has an excess risk in $O(d_{\pi}/n)$, as opposed to the standard $O(\sqrt{d_{\pi}/n})$, where $n$…

Machine Learning · Statistics 2025-03-03 Charles Riou , Pierre Alquier , Badr-Eddine Chérief-Abdellatif

In this paper, we consider the related problems of multicalibration -- a multigroup fairness notion and omniprediction -- a simultaneous loss minimization paradigm, both in the distributional and online settings. The recent work of Garg et…

Machine Learning · Computer Science 2025-05-29 Haipeng Luo , Spandan Senapati , Vatsal Sharan

While the SLIM approach obtained high ranking-accuracy in many experiments in the literature, it is also known for its high computational cost of learning its parameters from data. For this reason, we focus in this paper on variants of…

Information Retrieval · Computer Science 2019-05-01 Harald Steck

In stochastic optimization, the population risk is generally approximated by the empirical risk. However, in the large-scale setting, minimization of the empirical risk may be computationally restrictive. In this paper, we design an…

Machine Learning · Statistics 2016-11-22 Murat A. Erdogdu , Mohsen Bayati , Lee H. Dicker

In some applied scenarios, the availability of complete data is restricted, often due to privacy concerns; only aggregated, robust and inefficient statistics derived from the data are made accessible. These robust statistics are not…

Methodology · Statistics 2024-02-23 Antoine Luciano , Christian P. Robert , Robin J. Ryder

Sparse model estimation is a topic of high importance in modern data analysis due to the increasing availability of data sets with a large number of variables. Another common problem in applied statistics is the presence of outliers in the…

Applications · Statistics 2025-02-03 Andreas Alfons , Christophe Croux , Sarah Gelper

Almost all scientific data have uncertainties originating from different sources. Gaussian process regression (GPR) models are a natural way to model data with Gaussian-distributed uncertainties. GPR also has the benefit of reducing I/O…

Machine Learning · Statistics 2025-12-16 Haoyu Li , Isaac J Michaud , Ayan Biswas , Han-Wei Shen

The Gibbs sampler (GS) is a crucial algorithm for approximating complex calculations, and it is justified by Markov chain theory, the alternating projection theorem, and $I$-projection, separately. We explore the equivalence between these…

Computation · Statistics 2024-10-15 Kun-Lin Kuo , Yuchung J. Wang

Replicability requires that algorithmic conclusions remain consistent when rerun on independently drawn data. A central structural question is composition: given $k$ problems each admitting a $\rho$-replicable algorithm with sample…

‹ Prev 1 8 9 10 Next ›