English
Related papers

Related papers: A Partial EM Algorithm for Clustering White Breads

200 papers

LLM pre-training efficacy increasingly depends on data composition rather than sheer volume. Yet, optimal mixing is hindered by categorization flaws: human taxonomies suffer from ontological misalignment, and Euclidean clustering fails to…

Machine Learning · Computer Science 2026-05-27 Yue Min , Ziyun Qiao , Ruining Chen , Yujun Li

Influence Maximization (IM) aims to maximize the number of people that become aware of a product by finding the `best' set of `seed' users to initiate the product advertisement. Unlike prior arts on static social networks containing fixed…

Social and Information Networks · Computer Science 2019-11-14 Xudong Wu , Luoyi Fu , Zixin Zhang , Jingfan Meng , Xinbing Wang , Guihai Chen

The stochastic blockmodel (SBM) models the connectivity within and between disjoint subsets of nodes in networks. Prior work demonstrated that the rows of an SBM's adjacency spectral embedding (ASE) and Laplacian spectral embedding (LSE)…

Methodology · Statistics 2022-05-04 Zachary M. Pisano , Joshua S. Agterberg , Carey E. Priebe , Daniel Q. Naiman

Advanced metering infrastructure (AMI) enables utilities to obtain granular energy consumption data, which offers a unique opportunity to design customer segmentation strategies based on their impact on various operational metrics in…

Applications · Statistics 2020-03-13 Yuxuan Yuan , Kaveh Dehghanpour , Fankun Bu , Zhaoyu Wang

The Expectation-Maximization algorithm is perhaps the most broadly used algorithm for inference of latent variable problems. A theoretical understanding of its performance, however, largely remains lacking. Recent results established that…

Machine Learning · Statistics 2019-05-30 Jeongyeol Kwon , Wei Qian , Constantine Caramanis , Yudong Chen , Damek Davis

In this paper, different strands of literature are combined in order to obtain algorithms for semi-parametric estimation of discrete choice models that include the modelling of unobserved heterogeneity by using mixing distributions for the…

Methodology · Statistics 2022-12-12 Dietmar Bauer , Sebastian Büscher , Manuel Batram

Complex biological processes are usually experimented along time among a collection of individuals. Longitudinal data are then available and the statistical challenge is to better understand the underlying biological mechanisms. The…

Statistics Theory · Mathematics 2015-06-11 Pierre Barbillon , Célia Barthélémy , Adeline Samson

Building a shopping product collection has been primarily a human job. With the manual efforts of craftsmanship, experts collect related but diverse products with common shopping intent that are effective when displayed together, e.g.,…

Information Retrieval · Computer Science 2021-10-18 Hiun Kim , Jisu Jeong , Kyung-Min Kim , Dongjun Lee , Hyun Dong Lee , Dongpil Seo , Jeeseung Han , Dong Wook Park , Ji Ae Heo , Rak Yeong Kim

This paper proposes an uncertain data clustering approach to quantitatively analyze the complexity of prefabricated construction components through the integration of quality performance-based measures with associated engineering design…

Databases · Computer Science 2019-03-19 Wenying Ji , Simaan M. AbouRizk , Osmar R. Zaiane , Yitong Li

The classical mixture of linear experts (MoE) model is one of the widespread statistical frameworks for modeling, classification, and clustering of data. Built on the normality assumption of the error terms for mathematical and…

Methodology · Statistics 2020-07-15 Elham Mirfarah , Mehrdad Naderi , Ding-Geng Chen

The so-called matrix-element method (MEM) has long been used successfully as a classification tool in particle physics searches. In the presence of invisible final state particles, the traditional MEM typically assigns probabilities to an…

High Energy Physics - Phenomenology · Physics 2019-08-26 Stefan von Buddenbrock , Olivier Mattelaer , Michael Spannowsky

Product search is an important way for people to browse and purchase items on E-commerce platforms. While customers tend to make choices based on their personal tastes and preferences, analysis of commercial product search logs has shown…

Information Retrieval · Computer Science 2020-05-19 Keping Bi , Qingyao Ai , W. Bruce Croft

Studying competition and market structure at the product level instead of brand level can provide firms with insights on cannibalization and product line optimization. However, it is computationally challenging to analyze product-level…

Machine Learning · Computer Science 2020-05-22 Fanglin Chen , Xiao Liu , Davide Proserpio , Isamar Troncoso , Feiyu Xiong

Consider semi-supervised learning for classification, where both labeled and unlabeled data are available for training. The goal is to exploit both datasets to achieve higher prediction accuracy than just using labeled data alone. We…

Machine Learning · Statistics 2019-06-20 Xinwei Zhang , Zhiqiang Tan

We study partially linear models in settings where observations are arranged in independent groups but may exhibit within-group dependence. Existing approaches estimate linear model parameters through weighted least squares, with optimal…

Methodology · Statistics 2024-04-16 Elliot H. Young , Rajen D. Shah

We present a general method for fitting finite mixture models (FMM). Learning in a mixture model consists of finding the most likely cluster assignment for each data-point, as well as finding the parameters of the clusters themselves. In…

Machine Learning · Statistics 2019-12-20 Mathias Edman , Neil Dhir

Concept Bottleneck Model (CBM) is a methods for explaining neural networks. In CBM, concepts which correspond to reasons of outputs are inserted in the last intermediate layer as observed values. It is expected that we can interpret the…

Machine Learning · Statistics 2024-03-15 Naoki Hayashi , Yoshihide Sawada

Gaussian process is an indispensable tool in clustering functional data, owing to it's flexibility and inherent uncertainty quantification. However, when the functional data is observed over a large grid (say, of length $p$), Gaussian…

Computation · Statistics 2023-09-15 Anirban Chakraborty , Abhisek Chakraborty

This paper develops the likelihood ratio-based test of the null hypothesis of a M0-component model against an alternative of (M0 + 1)-component model in the normal mixture panel regression by extending the Expectation-Maximization (EM) test…

Econometrics · Economics 2023-06-06 Yu Hao , Hiroyuki Kasahara

In this work, we propose an original method for aggregating multiple clustering coming from different sources of information. Each partition is encoded by a co-membership matrix between observations. Our approach uses a mixture of…

Machine Learning · Computer Science 2024-01-10 Kylliann De Santiago , Marie Szafranski , Christophe Ambroise
‹ Prev 1 8 9 10 Next ›