English
Related papers

Related papers: The noise level in linear regression with dependen…

200 papers

We adapt arguments concerning information-theoretic convergence in the Central Limit Theorem to the case of dependent random variables under Rosenblatt mixing conditions. The key is to work with random variables perturbed by the addition of…

Probability · Mathematics 2008-10-06 Oliver Johnson

In this article, we introduce a conditional marginal model for longitudinal data, in which the residuals form a martingale difference sequence. This model allows us to consider a rich class of estimating equations, which contains several…

Statistics Theory · Mathematics 2008-07-15 R. M. Balan , L. Dumitrescu , I. Schiopu-Kratina

This paper studies the problem of shuffled linear regression, where the correspondence between predictors and responses in a linear model is obfuscated by a latent permutation. Specifically, we consider the model $y = \Pi_* X \beta_* + w$,…

Statistics Theory · Mathematics 2024-02-16 Leon Lufkin , Yihong Wu , Jiaming Xu

There is a growing need for models that are interpretable and have reduced energy and computational cost (e.g., in health care analytics and federated learning). Examples of algorithms to train such models include logistic regression and…

Machine Learning · Computer Science 2023-02-21 Tyler Sypherd , Nathan Stromberg , Richard Nock , Visar Berisha , Lalitha Sankar

We study the problem of high-dimensional linear regression in a robust model where an $\epsilon$-fraction of the samples can be adversarially corrupted. We focus on the fundamental setting where the covariates of the uncorrupted samples are…

Machine Learning · Computer Science 2018-06-04 Ilias Diakonikolas , Weihao Kong , Alistair Stewart

Is it possible to perform linear regression on datasets whose labels are shuffled with respect to the inputs? We explore this question by proposing several estimators that recover the weights of a noisy linear model from labels that are…

Machine Learning · Statistics 2017-05-05 Abubakar Abid , Ada Poon , James Zou

Annotation errors are a challenge not only during training of machine learning models, but also during their evaluation. Label variations and inaccuracies in datasets often manifest as contradictory examples that deviate from established…

Computer Vision and Pattern Recognition · Computer Science 2025-01-23 David Tschirschwitz , Volker Rodehorst

Uncertainty is ubiquitous in real-world data, and the assumptions underlying classical linear regression models are often violated in practice. Inspired by the theory of sublinear expectation, we consider a linear regression model where the…

Statistics Theory · Mathematics 2026-04-28 Xifeng Li , Shuzhen Yang

Data on rates, percentages or proportions arise frequently in many different applied disciplines like medical biology, health care, psychology and several others. In this paper, we develop a robust inference procedure for the beta…

Methodology · Statistics 2018-01-16 Abhik Ghosh

The phenomenon of benign overfitting, where a predictor perfectly fits noisy training data while attaining near-optimal expected loss, has received much attention in recent years, but still remains not fully understood beyond well-specified…

Machine Learning · Computer Science 2023-04-18 Ohad Shamir

Supervisory signals have the potential to make low-dimensional data representations, like those learned by mixture and topic models, more interpretable and useful. We propose a framework for training latent variable models that explicitly…

We establish optimal convergence rates up to a log-factor for a class of deep neural networks in a classification setting under a restraint sometimes referred to as the Tsybakov noise condition. We construct classifiers in a general setting…

Statistics Theory · Mathematics 2022-07-26 Joseph T. Meyer

We develop a maximum-likelihood based method for regression in a setting where the dependent variable is a random graph and covariates are available on a graph-level. The model generalizes the well-known $\beta$-model for random graphs by…

Methodology · Statistics 2017-05-24 Johan Wahlström , Isaac Skog , Patricio S. La Rosa , Peter Händel , Arye Nehorai

A meta-model of the input-output data of a computationally expensive simulation is often employed for prediction, optimization, or sensitivity analysis purposes. Fitting is enabled by a designed experiment, and for computationally expensive…

Methodology · Statistics 2023-12-01 Andrew Gill , David J. Warne , Antony M. Overstall , Clare McGrory , James M. McGree

Recently, various algorithms for data-driven simulation and control have been proposed based on the Willems' fundamental lemma. However, when collected data are noisy, these methods lead to ill-conditioned data-driven model structures. In…

Systems and Control · Electrical Eng. & Systems 2023-03-20 Mingzhou Yin , Andrea Iannelli , Roy S. Smith

We study the problem of learning mixtures of low-rank models, i.e. reconstructing multiple low-rank matrices from unlabelled linear measurements of each. This problem enriches two widely studied settings -- low-rank matrix sensing and mixed…

Machine Learning · Statistics 2021-03-10 Yanxi Chen , Cong Ma , H. Vincent Poor , Yuxin Chen

The ultimate goal of a supervised learning algorithm is to produce models constructed on the training data that can generalize well to new examples. In classification, functional margin maximization -- correctly classifying as many training…

Machine Learning · Computer Science 2020-01-29 Nikolaos Nikolaou , Henry Reeve , Gavin Brown

We develop novel empirical Bernstein inequalities for the variance of bounded random variables. Our inequalities hold under constant conditional variance and mean, without further assumptions like independence or identical distribution of…

Statistics Theory · Mathematics 2026-05-28 Diego Martinez-Taboada , Aaditya Ramdas

We study a special case of the problem of statistical learning without the i.i.d. assumption. Specifically, we suppose a learning method is presented with a sequence of data points, and required to make a prediction (e.g., a classification)…

Machine Learning · Computer Science 2018-05-22 Steve Hanneke , Liu Yang

Prediction via deterministic continuous-time models will always be subject to model error, for example due to unexplainable phenomena, uncertainties in any data driving the model, or discretisation/resolution issues. In this paper, we build…

Dynamical Systems · Mathematics 2025-06-30 Liam Blake , John Maclean , Sanjeeva Balasuriya