English
Related papers

Related papers: On unbalanced data and common shock models in stoc…

200 papers

Loss reserving generally focuses on identifying a single model that can generate superior predictive performance. However, different loss reserving models specialise in capturing different aspects of loss data. This is recognised in…

Methodology · Statistics 2024-06-04 Benjamin Avanzi , Yanfeng Li , Bernard Wong , Alan Xian

Studies of the effects of medical interventions increasingly take place in distributed research settings using data from multiple clinical data sources including electronic health records and administrative claims. In such settings, privacy…

Methodology · Statistics 2021-01-06 Martijn J. Schuemie , Yong Chen , David Madigan , Marc A. Suchard

Empirical Bayes methods are widely used for large-scale inference, yet most classical approaches assume homoscedastic observations and focus primarily on posterior mean estimation. We develop a nonparametric empirical Bayes framework for…

Methodology · Statistics 2026-04-24 Zhigen Zhao , Shonosuke Sugaasawa

The reliable operation of a power distribution system relies on a good prior knowledge of its topology and its system state. Although crucial, due to the lack of direct monitoring devices on the switch statuses, the topology information is…

Systems and Control · Electrical Eng. & Systems 2021-10-19 Yijun Xu , Jaber Valinejad , Mert Korkali , Lamine Mili , Yajun Wang , Xiao Chen , Zongsheng Zheng

Tweedie's formula is central to measurement-error analysis and empirical Bayes. Under Gaussian noise, the formula identifies the posterior mean directly from the observed-data density, bypassing nonparametric deconvolution. Beyond a few…

Statistics Theory · Mathematics 2026-05-05 Santiago Torres

The development of unsupervised hashing is advanced by the recent popular contrastive learning paradigm. However, previous contrastive learning-based works have been hampered by (1) insufficient data similarity mining based on global-only…

Computer Vision and Pattern Recognition · Computer Science 2022-09-29 Jiaguo Yu , Huming Qiu , Dubing Chen , Haofeng Zhang

Data aggregation, also known as meta analysis, is widely used to combine knowledge on parameters shared in common (e.g., average treatment effect) between multiple studies. In this paper, we introduce an attractive data aggregation scheme…

Methodology · Statistics 2023-05-10 Snigdha Panigrahi , Jingshen Wang , Xuming He

The sensitivity of loss reserving techniques to outliers in the data or deviations from model assumptions is a well known challenge. It has been shown that the popular chain-ladder reserving approach is at significant risk to such aberrant…

Methodology · Statistics 2023-06-22 Benjamin Avanzi , Mark Lavender , Greg Taylor , Bernard Wong

A new class of probabilistic models for cascading failure propagation in interconnected systems is proposed. The models take into account important characteristics of real systems that are not considered in existing generic approaches.…

Disordered Systems and Neural Networks · Physics 2010-03-31 Jörg Lehmann , Jakob Bernasconi

Motivated by two case studies using primary care records from the Clinical Practice Research Datalink, we describe statistical methods that facilitate the analysis of tall data, with very large numbers of observations. Our focus is on…

Methodology · Statistics 2018-05-14 Kirsty Rhodes , Rebecca Turner , Rupert Payne , Ian White

In clinical settings, we often face the challenge of building prediction models based on small observational data sets. For example, such a data set might be from a medical center in a multi-center study. Differences between centers might…

Balanced representation learning methods have been applied successfully to counterfactual inference from observational data. However, approaches that account for survival outcomes are relatively limited. Survival data are frequently…

Machine Learning · Statistics 2021-03-04 Paidamoyo Chapfuwa , Serge Assaad , Shuxi Zeng , Michael J. Pencina , Lawrence Carin , Ricardo Henao

Statistical inference for stochastic block models typically relies on the spectrum of the normalized adjacency matrix $\A^*$. In practice, the true probability matrix $\mathbf{B}$ is unknown and must be replaced by a plug-in estimator…

Methodology · Statistics 2026-04-09 Jianwei Hu , Ding Chen , Ji Zhu

Tweedie regression models provide a flexible family of distributions to deal with non-negative highly right-skewed data as well as symmetric and heavy tailed data and can handle continuous data with probability mass at zero. The estimation…

Methodology · Statistics 2017-04-25 Wagner H. Bonat , Célestin C. Kokonendji

Two-part models and Tweedie generalized linear models (GLMs) have been used to model loss costs for short-term insurance contract. For most portfolios of insurance claims, there is typically a large proportion of zero claims that leads to…

Applications · Statistics 2020-06-11 Zhiyu Quan , Zhiguo Wang , Guojun Gan , Emiliano A. Valdez

A composite loss framework is proposed for low-rank modeling of data consisting of interesting and common values, such as excess zeros or missing values. The methodology is motivated by the generalized low-rank framework and the hurdle…

Machine Learning · Statistics 2017-09-07 Christopher Dienes

Tweedie exponential dispersion family constitutes a fairly rich sub-class of the celebrated exponential family. In particular, a member, compound Poisson gamma (CP-g) model has seen extensive use over the past decade for modeling mixed…

Applications · Statistics 2023-02-14 Aritra Halder , Shariq Mohammed , Kun Chen , Dipak Dey

Multiple sets of measurements on the same objects obtained from different platforms may reflect partially complementary information of the studied system. The integrative analysis of such data sets not only provides us with the opportunity…

Methodology · Statistics 2020-10-15 Yipeng Song , Johan A. Westerhuis , Age K. Smilde

Statistical agencies and other institutions collect data under the promise to protect the confidentiality of respondents. When releasing microdata samples, the risk that records can be identified must be assessed. To this aim, a widely…

Applications · Statistics 2015-06-03 Cinzia Carota , Maurizio Filippone , Roberto Leombruni , Silvia Polettini

This paper studies a Markov network model for unbalanced data, aiming to solve the problems of classification bias and insufficient minority class recognition ability of traditional machine learning models in environments with uneven class…

Machine Learning · Computer Science 2025-02-06 Junliang Du , Shiyu Dou , Bohuan Yang , Jiacheng Hu , Tai An