中文
相关论文

相关论文: Adjusting the adjusted Rand Index -- A multinomial…

200 篇论文

Finite Mixture of Regressions (FMR) models are among the most widely used approaches in dealing with the heterogeneity among the observations in regression problems. One of the limitations of current approaches is their inability to…

应用统计 · 统计学 2018-06-25 Haidar Almohri , Arash Ali Amini , Ratna Babu Chinnam

We present a consensus Monte Carlo algorithm that scales existing Bayesian nonparametric models for clustering and feature allocation to big data. The algorithm is valid for any prior on random subsets such as partitions and latent feature…

统计计算 · 统计学 2020-02-26 Yang Ni , Yuan Ji , Peter Mueller

Clustered observations are ubiquitous in controlled and observational studies and arise naturally in multi-centre trials or longitudinal surveys. We present a novel model for the analysis of clustered observations where the marginal…

统计方法学 · 统计学 2022-11-03 Luisa Barbanti , Torsten Hothorn

It is common to conduct causal inference in matched observational studies by proceeding as though treatment assignments within matched sets are assigned uniformly at random and using this distribution as the basis for inference. This…

统计方法学 · 统计学 2023-11-14 Samuel D. Pimentel , Yaxuan Huang

It has become standard for empirical studies to conduct inference robust to cluster dependence and heterogeneity. With a small number of clusters, the normal approximation for the $t$-statistics of regression coefficients may be poor. This…

计量经济学 · 经济学 2026-03-27 Bulat Gafarov , Takuya Ura

Motivation: Clustering is a frequently used concept in variety of bioinformatical applications. We present a new method for hierarchical clustering of data called mutual information clustering (MIC) algorithm. It uses mutual information…

定量方法 · 定量生物学 2007-05-23 Alexander Kraskov , Harald Stögbauer , Ralph G. Andrzejak , Peter Grassberger

This report discusses two new indices for comparing clusterings of a set of points. The motivation for looking at new ways for comparing clusterings stems from the fact that the existing clustering indices are based on set cardinality alone…

机器学习 · 计算机科学 2014-12-01 Zaeem Hussain , Marina Meila

We propose Adaptive Incremental Mixture Markov chain Monte Carlo (AIMM), a novel approach to sample from challenging probability distributions defined on a general state-space. While adaptive MCMC methods usually update a parametric…

统计方法学 · 统计学 2018-06-01 Florian Maire , Nial Friel , Antonietta Mira , Adrian Raftery

The analysis of continously larger datasets is a task of major importance in a wide variety of scientific fields. In this sense, cluster analysis algorithms are a key element of exploratory data analysis, due to their easiness in the…

机器学习 · 统计学 2018-01-10 Marco Capó , Aritz Pérez , Jose A. Lozano

We propose a change-point detection method for large scale multiple testing problems with data having clustered signals. Unlike the classic change-point setup, the signals can vary in size within a cluster. The clustering structure on the…

统计方法学 · 统计学 2021-10-07 Hongyuan Cao , Wei Biao Wu

Whether class labels in a given data set correspond to meaningful clusters is crucial for the evaluation of clustering algorithms using real-world data sets. This property can be quantified by separability measures. The central aspects of…

机器学习 · 统计学 2025-04-11 Jana Gauss , Fabian Scheipl , Moritz Herrmann

Machine learning models often generalize poorly to out-of-distribution (OOD) data as a result of relying on features that are spuriously correlated with the label during training. Recently, the technique of Invariant Risk Minimization (IRM)…

机器学习 · 计算机科学 2023-01-18 Dongsung Huh , Avinash Baidya

Datasets composed of numerical and categorical attributes (also called mixed data hereinafter) are common in real clustering tasks. Differing from numerical attributes that indicate tendencies between two concepts (e.g., high and low…

机器学习 · 计算机科学 2026-03-06 Yiqun Zhang , Mingjie Zhao , Yizhou Chen , Yang Lu , Yiu-ming Cheung

Change point analysis has applications in a wide variety of fields. The general problem concerns the inference of a change in distribution for a set of time-ordered observations. Sequential detection is an online version in which new data…

统计方法学 · 统计学 2013-10-16 David S. Matteson , Nicholas A. James

Time series segmentation is a fundamental task in analyzing temporal data across various domains, from human activity recognition to energy monitoring. While numerous state-of-the-art methods have been developed to tackle this problem, the…

机器学习 · 计算机科学 2025-10-28 Félix Chavelli , Paul Boniol , Michaël Thomazo

Many scientific fields, including human gut microbiome science, collect multivariate count data where the sum of the counts is unrelated to the scale of the underlying system being measured (e.g., total microbial load in a subject's colon).…

Betweenness is a well-known centrality measure that ranks the nodes according to their participation in the shortest paths of a network. In several scenarios, having a high betweenness can have a positive impact on the node itself. Hence,…

Automatic data augmentation (AutoDA) plays an important role in enhancing the generalization of neural networks. However, mainstream AutoDA methods often encounter two challenges: either the search process is excessively time-consuming,…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Anqi Xiao , Weichen Yu , Hongyuan Yu

We study the sampling of spatial fields using sensors that are location-unaware but deployed according to a known statistical distribution. It has been shown that uniformly distributed location-unaware sensors cannot infer bandlimited…

信息论 · 计算机科学 2016-12-01 Ankur Mallick , Animesh Kumar

This article concerns the dimension reduction in regression for large data set. We introduce a new method based on the sliced inverse regression approach, called cluster-based regularized sliced inverse regression. Our method not only keeps…

应用统计 · 统计学 2013-12-03 Yue Yu , Zhihong Chen , Jie Yang