中文
相关论文

相关论文: Extension of the Dip-test Repertoire -- Efficient …

200 篇论文

Detecting multimodality in empirical distributions is a fundamental problem in statistics and data analysis, with applications ranging from clustering to the study of complex systems. In practice, however, assessing departures from…

统计方法学 · 统计学 2026-05-21 Edoardo Di Martino , Matteo Cinelli , Roy Cerqueti

This paper provides a new unimodality test with application in hierarchical clustering methods. The proposed method denoted by signature test (Sigtest), transforms the data based on its statistics. The transformed data has much smaller…

机器学习 · 计算机科学 2014-01-10 Mahdi Shahbaba , Soosan Beheshti

Many testing problems are readily amenable to randomised tests such as those employing data splitting. However despite their usefulness in principle, randomised tests have obvious drawbacks. Firstly, two analyses of the same dataset may…

统计方法学 · 统计学 2024-09-05 F. Richard Guo , Rajen D. Shah

Clustering is an essential data mining tool that aims to discover inherent cluster structure in data. As such, the study of clusterability, which evaluates whether data possesses such structure, is an integral part of cluster analysis. Yet,…

机器学习 · 计算机科学 2016-02-24 Margareta Ackerman , Andreas Adolfsson , Naomi Brownstein

Unimodality, pivotal in statistical analysis, offers insights into dataset structures and drives sophisticated analytical procedures. While unimodality's confirmation is straightforward for one-dimensional data using methods like…

统计方法学 · 统计学 2024-07-08 Prodromos Kolyvakis , Aristidis Likas

Cluster analysis is a fundamental research issue in statistics and machine learning. In many modern clustering methods, we need to determine whether two subsets of samples come from the same cluster. Since these subsets are usually…

机器学习 · 计算机科学 2025-07-15 Xinying Liu , Lianyu Hu , Mudi Jiang , Simeng Zhang , Jun Lou , Zengyou He

In this work, we introduce a novel methodology for divisive hierarchical clustering. Our divisive (``top-down'') approach is motivated by the fact that agglomerative hierarchical clustering (``bottom-up''), which is commonly used for…

统计方法学 · 统计学 2025-10-07 Jan O. Bauer

We are concerned with testing replicability hypotheses for many endpoints simultaneously. This constitutes a multiple test problem with composite null hypotheses. Traditional $p$-values, which are computed under least favourable parameter…

统计方法学 · 统计学 2020-02-26 Anh-Tuan Hoang , Thorsten Dickhaus

Selective inference is a subfield of statistics that enables valid inference after selection of a data-dependent question. In this paper, we introduce selectively dominant p-values, a class of p-values that allow practitioners to easily…

统计方法学 · 统计学 2024-11-22 Anav Sood

Many multiple testing procedures make use of the p-values from the individual pairs of hypothesis tests, and are valid if the p-value statistics are independent and uniformly distributed under the null hypotheses. However, it has recently…

统计方法学 · 统计学 2011-08-25 Joshua D. Habiger , Edsel A. Pena

The randomized $p$-value, (nonrandomized) mid-$p$-value and abstract randomized $p$-value have all been recommended for testing a null hypothesis whenever the test statistic has a discrete distribution. This paper provides a unifying…

统计计算 · 统计学 2014-12-02 Joshua D Habiger

Large-scale datasets have been pivotal to the advancements of deep learning models in recent years, but training on such large datasets invariably incurs substantial storage and computational overhead. Meanwhile, real-world datasets often…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Suorong Yang , Peng Ye , Wanli Ouyang , Dongzhan Zhou , Furao Shen

For many applications, it is critical to interpret and validate groups of observations obtained via clustering. A common validation approach involves testing differences in feature means between observations in two estimated clusters. In…

统计方法学 · 统计学 2023-11-29 Yiqun T. Chen , Lucy L. Gao

Dimension reduction and visualization of high-dimensional data have become very important research topics because of the rapid growth of large databases in data science. In this paper, we propose using a generalized sigmoid function to…

机器学习 · 统计学 2020-07-20 Yu Liang , Arin Chaudhuri , Haoyu Wang

Unsupervised learning has gained prominence in the big data era, offering a means to extract valuable insights from unlabeled datasets. Deep clustering has emerged as an important unsupervised category, aiming to exploit the non-linear…

机器学习 · 计算机科学 2024-02-02 Georgios Vardakas , Ioannis Papakostas , Aristidis Likas

We develop a greedy algorithm that is fast and scalable in the detection of a nested partition extracted from a dendrogram obtained from hierarchical clustering of a multivariate series. Our algorithm provides a $p$-value for each clade…

基因组学 · 定量生物学 2022-01-21 Christian Bongiorno , Salvatore Miccichè , Rosario N. Mantegna

The design of a metric between probability distributions is a longstanding problem motivated by numerous applications in Machine Learning. Focusing on continuous probability distributions on the Euclidean space $\mathbb{R}^d$, we introduce…

$P$-values that are derived from continuously distributed test statistics are typically uniformly distributed on $(0,1)$ under least favorable parameter configurations (LFCs) in the null hypothesis. Conservativeness of a $p$-value $P$…

统计方法学 · 统计学 2023-03-13 Daniel Ochieng , Anh-Tuan Hoang , Thorsten Dickhaus

This paper introduces an open-ended sequential algorithm for computing the p-value of a test using Monte Carlo simulation. It guarantees that the resampling risk, the probability of a different decision than the one based on the theoretical…

统计理论 · 数学 2013-07-30 Axel Gandy

We consider the problem of testing for a difference in means between clusters of observations identified via k-means clustering. In this setting, classical hypothesis tests lead to an inflated Type I error rate. To overcome this problem, we…

统计方法学 · 统计学 2022-03-30 Yiqun T. Chen , Daniela M. Witten
‹ 上一页 1 2 3 10 下一页 ›