English
Related papers

Related papers: Combining Structural and Unstructured Data: A Topi…

200 papers

In this article, we propose two classes of semiparametric mixture regression models with single-index for model based clustering. Unlike many semiparametric/nonparametric mixture regression models that can only be applied to low dimensional…

Methodology · Statistics 2017-08-15 Sijia Xiang , Weixin Yao

Traditionally, the detection of fraudulent insurance claims relies on business rules and expert judgement which makes it a time-consuming and expensive process (\'Oskarsd\'ottir et al., 2022). Consequently, researchers have been examining…

Machine Learning · Computer Science 2024-10-08 Bavo D. C. Campo , Katrien Antonio

Mixtures of linear mixed models (MLMMs) are useful for clustering grouped data and can be estimated by likelihood maximization through the EM algorithm. The conventional approach to determining a suitable number of components is to compare…

Applications · Statistics 2014-05-26 Siew Li Tan , David J. Nott

In a context of constant increase in competition and heightened regulatory pressure, accuracy, actuarial precision, as well as transparency and understanding of the tariff, are key issues in non-life insurance. Traditionally used…

Applications · Statistics 2025-03-28 Markéta Krùpovà , Nabil Rachdi , Quentin Guibert

This article describes posterior maximization for topic models, identifying computational and conceptual gains from inference under a non-standard parametrization. We then show that fitted parameters can be used as the basis for a novel…

Applications · Statistics 2011-12-30 Matthew A. Taddy

We consider the estimation of Dirichlet Process Mixture Models (DPMMs) in distributed environments, where data are distributed across multiple computing nodes. A key advantage of Bayesian nonparametric models such as DPMMs is that they…

Machine Learning · Statistics 2017-09-20 Ruohui Wang , Dahua Lin

Topic models have been the prominent tools for automatic topic discovery from text corpora. Despite their effectiveness, topic models suffer from several limitations including the inability of modeling word ordering information in…

Computation and Language · Computer Science 2022-02-10 Yu Meng , Yunyi Zhang , Jiaxin Huang , Yu Zhang , Jiawei Han

Multi-component polymer systems are of interest in organic photovoltaic and drug delivery applications, among others where diverse morphologies influence performance. An improved understanding of morphology classification, driven by…

Computational Engineering, Finance, and Science · Computer Science 2020-08-27 Pavan Inguva , Lachlan Mason , Indranil Pan , Miselle Hengardi , Omar K. Matar

Item nonresponse is frequently encountered in practice. Ignoring missing data can lose efficiency and lead to misleading inference. Fractional imputation is a frequentist approach of imputation for handling missing data. However, the…

Methodology · Statistics 2018-09-18 Hejian Sang , Jae Kwang Kim

The paper proposes a latent variable model for binary data coming from an unobserved heterogeneous population. The heterogeneity is taken into account by replacing the traditional assumption of Gaussian distributed factors by a finite…

Methodology · Statistics 2010-10-13 Silvia Cagnone , Cinzia Viroli

This paper addresses the identification of insurance models with multidimensional screening where insurees have private information about their risk and risk aversion. The model includes a random damage and the possibility of several…

Mathematical Finance · Quantitative Finance 2016-01-18 Gaurab Aryal , Isabelle Perrigne , Quang Vuong

Accurate loss reserving is crucial in Property and Casualty (P&C) insurance for financial stability, regulatory compliance, and effective risk management. We propose a novel micro-level Cox model based on hidden Markov models (HMMs).…

Applications · Statistics 2026-01-14 Hassan Abdelrahman , Andrei Badescu , Radu Craiu , Sheldon Lin

We are concerned in clustering continuous data sets subject to non-ignorable missingness. We perform clustering with a specific semi-parametric mixture, under the assumption of conditional independence given the component. The mixture model…

Methodology · Statistics 2021-07-20 Marie Du Roy de Chaumaray , Matthieu Marbac

Our paper explores a discrete-time risk model with time-varying premiums, investigating two types of correlated claims: main claims and by-claims. Settlement of the by-claims can be delayed for one time period, representing real-world…

Risk Management · Quantitative Finance 2024-08-02 Dhiti Osatakul , Shuanming Li , Xueyuan Wu

Reinsurance optimization is a cornerstone of solvency and capital management, yet traditional approaches often rely on restrictive distributional assumptions and static program designs. We propose a hybrid framework that combines…

Econometrics · Economics 2026-03-24 Stella C. Dong

Imputing missing values is an important preprocessing step in data analysis, but the literature offers little guidance on how to choose between different imputation models. This letter suggests adopting the imputation model that generates a…

Methodology · Statistics 2021-07-13 Moritz Marbach

In this paper, we propose two important extensions to cluster-weighted models (CWMs). First, we extend CWMs to have generalized cluster-weighted models (GCWMs) by allowing modeling of non-Gaussian distribution of the continuous covariates,…

Applications · Statistics 2019-01-01 Nikola Pocuca , Petar Jevtic , Paul D. McNicholas , Tatjana Miljkovic

The expectation-maximization (EM) algorithm and its variants are widely used in statistics. In high-dimensional mixture linear regression, the model is assumed to be a finite mixture of linear regression and the number of predictors is much…

Statistics Theory · Mathematics 2023-07-24 Ning Wang , Xin Zhang , Qing Mai

This study introduces a predictive maintenance strategy for high pressure industrial compressors using sensor data and features derived from unsupervised clustering integrated into classification models. The goal is to enhance model…

Machine Learning · Computer Science 2024-11-22 Alessandro Costa , Emilio Mastriani , Federico Incardona , Kevin Munari , Sebastiano Spinello

This PhD Thesis presents an investigation into the analysis of financial returns using mixture models, focusing on mixtures of generalized normal distributions (MGND) and their extensions. The study addresses several critical issues…

Statistical Finance · Quantitative Finance 2024-11-20 Pierdomenico Duttilo