English
Related papers

Related papers: Zero-Inflated Logistic Regression Models with Shar…

200 papers

Additive two-tower models are popular learning-to-rank methods for handling biased user feedback in industry settings. Recent studies, however, report a concerning phenomenon: training two-tower models on clicks collected by well-performing…

Information Retrieval · Computer Science 2025-09-01 Philipp Hager , Onno Zoeter , Maarten de Rijke

The disaggregated time-series for the Consumer Price Index (CPI) often exhibits exact zero price changes, stemming from structural features of the data collection process. However, the currently prominent stochastic volatility model of…

Methodology · Statistics 2026-03-04 Geonhee Han , Kaoru Irie

Building classification models that predict a binary class label on the basis of high dimensional multi-omics datasets poses several challenges, due to the typically widely differing characteristics of the data layers in terms of number of…

Methodology · Statistics 2020-08-04 Alessandra Cabassi , Denis Seyres , Mattia Frontini , Paul D. W. Kirk

Logistic regression is a classical model for describing the probabilistic dependence of binary responses to multivariate covariates. We consider the predictive performance of the maximum likelihood estimator (MLE) for logistic regression,…

Statistics Theory · Mathematics 2026-02-20 Hugo Chardon , Matthieu Lerasle , Jaouad Mourtada

Microbiome `omics approaches can reveal intriguing relationships between the human microbiome and certain disease states. Along with the identification of specific bacteria taxa associated with diseases, recent scientific advancements…

Applications · Statistics 2019-10-07 Shuang Jiang , Guanghua Xiao , Andrew Y. Koh , Qiwei Li , Xiaowei Zhan

A new distribution on (0, 1), generalized Log-Lindley distribution, is proposed by extending the Log-Lindley distribution. This new distribution is shown to be a weighted Log-Lindley distribution. Important probabilistic and statistical…

Statistics Theory · Mathematics 2020-02-07 S. Chakraborty , S. H. Ong , C. M. Ng

Complementary-label learning is a weakly supervised learning problem in which each training example is associated with one or multiple complementary labels indicating the classes to which it does not belong. Existing consistent approaches…

Machine Learning · Computer Science 2024-10-14 Wei Wang , Takashi Ishida , Yu-Jie Zhang , Gang Niu , Masashi Sugiyama

Spin glass models, such as the Sherrington-Kirkpatrick, Hopfield and Ising models, are all well-studied members of the exponential family of discrete distributions, and have been influential in a number of application domains where they are…

Machine Learning · Statistics 2020-03-19 Constantinos Daskalakis , Nishanth Dikkala , Ioannis Panageas

Mixed-effects logistic regression is widely used for binary outcomes in hierarchical data, yet formal goodness-of-fit tests remain limited to random-intercept models and do not address sparse cluster settings. We extend a grouping-based…

Methodology · Statistics 2026-04-22 Ariel Linden

In subgroup analysis, testing the existence of a subgroup with a differential treatment effect serves as protection against spurious subgroup discovery. Despite its importance, this hypothesis testing possesses a complicated nature:…

Statistics Theory · Mathematics 2025-03-21 Shota Takeishi

The predominance of machine learning models in many spheres of human activity has led to a growing demand for their transparency. The transparency of models makes it possible to discern some factors, such as security or non-discrimination.…

Machine Learning · Computer Science 2026-01-16 Niffa Cheick Oumar Diaby , Thierry Duchesne , Mario Marchand

In reliability and life data analysis, the Weibull distribution is widely used to accommodate more data characteristics by changing the values of the parameters. We frequently observe many zeros or close to zero data points in reliability…

Methodology · Statistics 2022-06-06 Sumangal Bhattacharya , Ishapathik Das , Muralidharan Kunnummal

Binomial data with unknown sizes often appear in biological and medical sciences. The previous methods either use the Poisson approximation or the quasi-likelihood approach. A full likelihood approach is proposed by treating unknown sizes…

Statistics Theory · Mathematics 2007-06-13 Wei Zhang

Reliably characterizing the full conditional distribution of a multivariate response variable given a set of covariates is crucial for trustworthy decision-making. However, misspecified or miscalibrated multivariate models may yield a poor…

Machine Learning · Computer Science 2025-10-27 Victor Dheur , Souhaib Ben Taieb

The literature has covered the features and uses of the traditional univariate and bivariate logistic distributions in great detail. It is reasonable to wonder, though, if logistic marginals and conditionals could exhibit a similar…

Applications · Statistics 2023-11-15 Banoth Veeranna

In many cases it makes sense to model a relationship symmetrically, not implying any particular directionality. Consider the classical example of a recommendation system where the rating of an item by a user should symmetrically be…

Artificial Intelligence · Computer Science 2012-07-02 Zhao Xu , Volker Tresp , Kai Yu , Hans-Peter Kriegel

A load sharing system has several components and the failure of one component can affect the lifetime of the surviving components. Since component failure does not equate to system failure for different system designs, the analysis of the…

Applications · Statistics 2023-07-20 Tim Pesch , Erhard Cramer , Edward Cripps , Adriano Polpo

Finite mixtures are a flexible modeling tool for irregularly shaped densities and samples from heterogeneous populations. When modeling with mixtures using an exchangeable prior on the component features, the component labels are arbitrary…

Methodology · Statistics 2020-07-10 Deborah Kunkel , Mario Peruggia

In regression models with missing outcomes, selection bias can arise when the missingness mechanism depends on the outcome itself. This proposal focuses on an extension of the Heckman model to a setting where the outcome is binary and both…

Methodology · Statistics 2025-11-18 Marco Doretti , Elena Stanghellini , Alessandro Taraborrelli

Data privacy and security becomes a major concern in building machine learning models from different data providers. Federated learning shows promise by leaving data at providers locally and exchanging encrypted information. This paper…

Machine Learning · Computer Science 2019-12-05 Kai Yang , Tao Fan , Tianjian Chen , Yuanming Shi , Qiang Yang