English
Related papers

Related papers: A robust contaminated discrete Weibull regression …

200 papers

For many machine learning algorithms, two main assumptions are required to guarantee performance. One is that the test data are drawn from the same distribution as the training data, and the other is that the model is correctly specified.…

Machine Learning · Computer Science 2020-02-03 Kun Kuang , Ruoxuan Xiong , Peng Cui , Susan Athey , Bo Li

Distributionally robust optimization (DRO) is an effective approach for data-driven decision-making in the presence of uncertainty. Geometric uncertainty due to sampling or localized perturbations of data points is captured by Wasserstein…

Machine Learning · Statistics 2023-11-10 Sloan Nietert , Ziv Goldfeld , Soroosh Shafiee

In this paper, we develop connections between two seemingly disparate, but central, models in robust statistics: Huber's epsilon-contamination model and the heavy-tailed noise model. We provide conditions under which this connection…

Machine Learning · Statistics 2019-07-03 Adarsh Prasad , Sivaraman Balakrishnan , Pradeep Ravikumar

The integration of external data using Bayesian mixture priors has become a powerful approach in clinical trials, offering significant potential to improve trial efficiency. Despite their strengths in analytical tractability and practical…

Methodology · Statistics 2025-10-07 Shouhao Zhou , Qiuxin Gao , Chenqi Fu , Yanxun Xu

One of the key challenges in predictive maintenance is to predict the impending downtime of an equipment with a reasonable prediction horizon so that countermeasures can be put in place. Classically, this problem has been posed in two…

Machine Learning · Computer Science 2018-12-19 Karan Aggarwal , Onur Atan , Ahmed Farahat , Chi Zhang , Kosta Ristovski , Chetan Gupta

Out-of-distribution (OOD) detection is an important task in machine learning systems for ensuring their reliability and safety. Deep probabilistic generative models facilitate OOD detection by estimating the likelihood of a data sample.…

Machine Learning · Computer Science 2021-06-16 Jaemoo Choi , Changyeon Yoon , Jeongwoo Bae , Myungjoo Kang

Deep neural network (DNN) regression models are widely used in applications requiring state-of-the-art predictive accuracy. However, until recently there has been little work on accurate uncertainty quantification for predictions from such…

Methodology · Statistics 2020-09-07 Nadja Klein , David J. Nott , Michael Stanley Smith

Importance weighting (IW) is a golden solver for joint distribution shift, where the joint distributions differ between the training and test data. To solve this problem, IW estimates test-to-training density ratios as importance weights…

Machine Learning · Computer Science 2026-05-26 Tongtong Fang , Nan Lu , Gang Niu , Kenji Fukumizu , Masashi Sugiyama

Recent advancements in diffusion models have demonstrated significant success in unsupervised anomaly segmentation. For anomaly segmentation, these models are first trained on normal data; then, an anomalous image is noised to an…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Mehrdad Moradi , Kamran Paynabar

Medical data often exhibit characteristics that make cluster analysis particularly challenging, such as missing values, outliers, and cluster features like skewness. Typically, such data would need to be preprocessed -- by cleaning outliers…

Methodology · Statistics 2025-12-16 Jason Pillay , Cristina Tortora , Antonio Punzo , Andriette Bekker

This paper presents a new methodology for generating continuous statistical distributions, integrating the exponentiated odds ratio within the framework of survival analysis. This new method enhances the flexibility and adaptability of…

Statistics Theory · Mathematics 2024-02-28 Xinyu Chen , Yuanqi Xie , Achraf Cohen , Shusen Pu

To ensure robust and reliable classification results, OoD (out-of-distribution) indicators based on deep generative models are proposed recently and are shown to work well on small datasets. In this paper, we conduct the first large…

Machine Learning · Computer Science 2021-08-13 Wenxiao Chen , Xiaohui Nie , Mingliang Li , Dan Pei

In data analysis, contamination caused by outliers is inevitable, and robust statistical methods are strongly demanded. In this paper, our concern is to develop a new approach for robust data analysis based on scoring rules. The scoring…

Statistics Theory · Mathematics 2013-11-22 Takafumi Kanamori , Hironori Fujisawa

We are concerned with the flexible parametric analysis of bivariate survival data. Elsewhere, we have extolled the virtues of the "power generalized Weibull" (PGW) distribution as an attractive vehicle for univariate parametric survival…

Methodology · Statistics 2019-01-11 M. C. Jones , Angela Noufaily , Kevin Burke

While the hurdle Poisson regression is a popular class of models for count data with excessive zeros, the link function in the binary component may be unsuitable for highly imbalanced cases. Ordinary Poisson regression is unable to handle…

Applications · Statistics 2020-08-14 Shuang Yin , Dipak K. Dey , Emiliano A. Valdez , Xiaomeng Li

Diffusion models achieve state-of-the-art generation quality across many applications, but their ability to capture rare or extreme events in heavy-tailed distributions remains unclear. In this work, we show that traditional diffusion and…

Machine Learning · Computer Science 2024-10-30 Kushagra Pandey , Jaideep Pathak , Yilun Xu , Stephan Mandt , Michael Pritchard , Arash Vahdat , Morteza Mardani

In various applications of heavy-tail modelling, the assumed Pareto behavior is tempered ultimately in the range of the largest data. In insurance applications, claim payments are influenced by claim management and claims may for instance…

Statistics Theory · Mathematics 2020-09-29 Jose Carlos Araujo Acuna , Hansjoerg Albrecher , Jan Beirlant

The Dirichlet-multinomial (DM) distribution plays a fundamental role in modern statistical methodology development and application. Recently, the DM distribution and its variants have been used extensively to model multivariate count data…

Methodology · Statistics 2023-02-27 Matthew D. Koslovsky

Recent work has shown that standard training via empirical risk minimization (ERM) can produce models that achieve high accuracy on average but low accuracy on underrepresented groups due to the prevalence of spurious features. A…

Machine Learning · Computer Science 2023-05-11 Yachuan Liu , Bohan Zhang , Qiaozhu Mei , Paramveer Dhillon

This article is aimed at the investigation of some properties of the Weibull cumulative exposure model on multiple-step step-stress accelerated life test data. Although the model includes a probabilistic idea of Miner's rule in order to…

Statistics Theory · Mathematics 2012-10-23 Yoshio Komori