English
Related papers

Related papers: Type I multivariate P\'olya-Aeppli distributions w…

200 papers

Over the years, data have become increasingly higher dimensional, which has prompted an increased need for dimension reduction techniques. This is perhaps especially true for clustering (unsupervised classification) as well as…

Methodology · Statistics 2019-11-21 Michael P. B. Gallaugher , Paul D. McNicholas

Testing and characterizing the difference between two data samples is of fundamental interest in statistics. Existing methods such as Kolmogorov-Smirnov and Cramer-von-Mises tests do not scale well as the dimensionality increases and…

Methodology · Statistics 2011-03-23 Li Ma , Wing H. Wong

When solving forecasting problems including multiple time-series features, existing approaches often fall into two extreme categories, depending on whether to utilize inter-feature information: univariate and complete-multivariate models.…

Artificial Intelligence · Computer Science 2024-08-20 Jaehoon Lee , Hankook Lee , Sungik Choi , Sungjun Cho , Moontae Lee

Many data sets cannot be accurately described by standard probability distributions due to the excess number of zero values present. For example, zero-inflation is prevalent in microbiome data and single-cell RNA sequencing data, which…

Methodology · Statistics 2024-11-20 Max Beveridge , Zach Goldstein , Hee Cheol Chung

A novel over-dispersed discrete distribution, namely the PoiTG distribution is derived by the convolution of a Poisson variate and an independently distributed transmuted geometric random variable. This distribution generalizes the…

Statistics Theory · Mathematics 2024-08-02 Anupama Nandi , Subrata Chakraborty , Aniket Biswas

Many multiple testing procedures make use of the p-values from the individual pairs of hypothesis tests, and are valid if the p-value statistics are independent and uniformly distributed under the null hypotheses. However, it has recently…

Methodology · Statistics 2011-08-25 Joshua D. Habiger , Edsel A. Pena

Binomial data with unknown sizes often appear in biological and medical sciences and are usually overdispersed. All previous methods used parametric models and only considered overdispersion due to the variation of sizes. The proposed…

Statistics Theory · Mathematics 2007-06-13 Wei Zhang

A family of parsimonious shifted asymmetric Laplace mixture models is introduced. We extend the mixture of factor analyzers model to the shifted asymmetric Laplace distribution. Imposing constraints on the constitute parts of the resulting…

Methodology · Statistics 2013-11-05 Brian C. Franczak , Paul D. McNicholas , Ryan P. Browne , Paula M. Murray

We propose a flexible model for count time series which has potential uses for both underdispersed and overdispersed data. The model is based on the Conway-Maxwell-Poisson (COM-Poisson) distribution with parameters varying along time to…

Computation · Statistics 2019-01-23 Ricardo S Ehlers

Diffusion models offer stable training and state-of-the-art performance for deep generative modeling tasks. Here, we consider their use in the context of multivariate subsurface modeling and probabilistic inversion. We first demonstrate…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Roberto Miele , Niklas Linde

Multivariate count data are defined as the number of items of different categories issued from sampling within a population, which individuals are grouped into categories. The analysis of multivariate count data is a recurrent and crucial…

Machine Learning · Statistics 2013-12-17 Pierre Fernique , Jean-Baptiste Durand , Yann Guédon

Although the specification of bivariate probability models using a collection of assumed conditional distributions is not a novel concept, it has received considerable attention in the last decade. In this study, a bivariate…

Methodology · Statistics 2025-03-20 Indranil Ghosh , Mina Norouzirad , Filipe J. Marques

We study the problem of multiple hypothesis testing for multidimensional data when inter-correlations are present. The problem of multiple comparisons is common in many applications. When the data is multivariate and correlated, existing…

Statistics Theory · Mathematics 2015-06-02 Mahdis Azadbakhsh , Xin Gao , Hanna Jankowski

We consider statistical procedures for hypothesis testing of real valued functionals of matched pairs with missing values. In order to improve the accuracy of existing methods, we propose a novel multiplication combination procedure.…

Statistics Theory · Mathematics 2018-01-29 Lubna Amro , Frank Konietschke , Markus Pauly

Claim frequency data in insurance records the number of claims on insurance policies during a finite period of time. Given that insurance companies operate with multiple lines of insurance business where the claim frequencies on different…

Applications · Statistics 2022-12-05 Pengcheng Zhang , David Pitt , Xueyuan Wu

We consider deep multivariate models for heterogeneous collections of random variables. In the context of computer vision, such collections may e.g. consist of images, segmentations, image attributes, and latent variables. When developing…

Machine Learning · Computer Science 2026-02-03 Dmitrij Schlesinger , Boris Flach , Alexander Shekhovtsov

In recent years, addressing the challenges posed by massive datasets has led researchers to explore aggregated data, particularly leveraging interval-valued data, akin to traditional symbolic data analysis. While much recent research, with…

Methodology · Statistics 2024-05-13 Ali Sadeghkhani , Abdolnasser Sadeghkhani

Copulas, generalized estimating equations, and generalized linear mixed models promote the analysis of grouped data where non-normal responses are correlated. Unfortunately, parameter estimation remains challenging in these three…

Methodology · Statistics 2024-10-16 Sarah S. Ji , Benjamin B. Chu , Hua Zhou , Kenneth Lange

It is usual to rely on the quasi-likelihood methods for deriving statistical methods applied to clustered multinomial data with no underlying distribution. Even though extensive literature can be encountered for these kind of data sets,…

Methodology · Statistics 2015-10-21 Juana María Alonso , Nirian Martín , Leandro Pardo

Event counts are response variables with non-negative integer values representing the number of times that an event occurs within a fixed domain such as a time interval, a geographical area or a cell of a contingency table. Analysis of…