English
Related papers

Related papers: On unbalanced data and common shock models in stoc…

200 papers

When random effects are correlated with sample design variables, the usual approach of employing individual survey weights (constructed to be inversely proportional to the unit survey inclusion probabilities) to form a pseudo-likelihood no…

Methodology · Statistics 2021-08-26 Terrance D. Savitsky , Matthew R. Williams

Recent advances in quantized compressed sensing and high-dimensional estimation have shown that signal recovery is even feasible under strong non-linear distortions in the observation process. An important characteristic of associated…

Information Theory · Computer Science 2023-08-08 Martin Genzel , Alexander Stollenwerk

Most research designing novel predictive models, or employing existing ones, assumes that training and testing data are independent and identically distributed. In practice, the data encountered at serving time often deviate from the…

Machine Learning · Computer Science 2026-03-30 Hanyu Duan , Yi Yang , Ahmed Abbasi , Kar Yan Tam

Data imbalance exists ubiquitously in real-world visual regressions, e.g., age estimation and pose estimation, hurting the model's generalizability and fairness. Thus, imbalanced regression gains increasing research attention recently.…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Jiawei Ren , Mingyuan Zhang , Cunjun Yu , Ziwei Liu

We introduce a new class of Poisson-exponential-Tweedie (PET) mixture in the framework of generalized linear models for ultra-overdispersed count data. The mean-variance relationship is of the form $m+m^{2}+\phi m^{p}$, where $\phi$ and $p$…

Methodology · Statistics 2019-08-26 Rahma Abid , Celestin C. Kokonendji , Afif Masmoudi

A new averaging method linking discrete to continuum variables of granular materials is developed and used to derive average balance equations. Its novelty lies in the choice of the decomposition between mean values and fluctuations of…

Soft Condensed Matter · Physics 2015-06-24 Riccardo Artoni , Patrick Richard

Root-mean-square error (RMSE) remains the default training loss for data-driven precipitation models, despite precipitation being semi-continuous, zero-inflated, strictly non-negative, and heavy-tailed. This Gaussian-implied objective…

Atmospheric and Oceanic Physics · Physics 2025-09-11 Kieran M. R. Hunt

Several problems in statistics involve the combination of high-variance unbiased estimators with low-variance estimators that are only unbiased under strong assumptions. A notable example is the estimation of causal effects while combining…

Methodology · Statistics 2023-05-25 Michael Oberst , Alexander D'Amour , Minmin Chen , Yuyan Wang , David Sontag , Steve Yadlowsky

Data imbalance remains one of the factors negatively affecting the performance of contemporary machine learning algorithms. One of the most common approaches to reducing the negative impact of data imbalance is preprocessing the original…

Machine Learning · Computer Science 2021-04-20 Michał Koziarski

In regression analysis of multivariate data, it is tacitly assumed that response and predictor variables in each observed response-predictor pair correspond to the same entity or unit. In this paper, we consider the situation of "permuted…

Statistics Theory · Mathematics 2017-11-17 Martin Slawski , Emanuel Ben-David

We consider joint inversion for two or more unknown parameters from observational data in the Bayesian framework. Standard approaches often either treat the parameters as independent or impose structural similarity through regularisation…

Methodology · Statistics 2026-05-04 Ruanui Nicholson , Matti Niskanen , Oliver J. Maclaren , Jari P. Kaipio

A relevant issue in panel data estimation is heteroscedasticity, which often occurs when the sample is large and individual units are of varying size. Furthermore, many of the available panel data sets are unbalanced in nature, because of…

Methodology · Statistics 2017-08-08 Silvia Platoni , Laura Barbieri , Daniele Moro , Paolo Sckokai

Data sharing barriers are paramount challenges arising from multicenter clinical trials where multiple data sources are stored in a distributed fashion at different local study sites. Merging such data sources into a common data storage for…

Methodology · Statistics 2022-04-05 Mengtong Hu , Xu Shi , Peter X. -K. Song

In this article, we present a model for analyzing the cooccurrence count data derived from practical fields such as user-item or item-item data from online shopping platform, cooccurring word-word pairs in sequences of texts. Such data…

Machine Learning · Computer Science 2025-01-03 Taejoon Kim , Haiyan Wang

It is common to conduct causal inference in matched observational studies by proceeding as though treatment assignments within matched sets are assigned uniformly at random and using this distribution as the basis for inference. This…

Methodology · Statistics 2023-11-14 Samuel D. Pimentel , Yaxuan Huang

It has been proposed that complex populations, such as those that arise in genomics studies, may exhibit dependencies among observations as well as among variables. This gives rise to the challenging problem of analyzing unreplicated…

Machine Learning · Statistics 2018-06-08 Michael Hornstein , Roger Fan , Kerby Shedden , Shuheng Zhou

We explore a statistical distribution that can simultaneously model the probability of zero outcome for non-users of health care utilization and continuous costs for users. We compare this distribution to other com- monly used models on…

Applications · Statistics 2019-01-09 Christoph Kurz

We study the estimation capacity of the generalized Lasso, i.e., least squares minimization combined with a (convex) structural constraint. While Lasso-type estimators were originally designed for noisy linear regression problems, it has…

Statistics Theory · Mathematics 2019-09-12 Martin Genzel , Gitta Kutyniok

If the prior probability distributions of all possible hypothetical true means and all possible observed means of a continuous variable are conditional on the universal set of all numbers (i.e., before the nature of a study is known and a…

Methodology · Statistics 2025-06-05 Huw Llewelyn

Recent empirical works have successfully used unlabeled data to learn feature representations that are broadly useful in downstream classification tasks. Several of these methods are reminiscent of the well-known word2vec embedding…

Machine Learning · Computer Science 2019-02-26 Sanjeev Arora , Hrishikesh Khandeparkar , Mikhail Khodak , Orestis Plevrakis , Nikunj Saunshi