中文
相关论文

相关论文: Adjusting the adjusted Rand Index -- A multinomial…

200 篇论文

We consider one of the most basic multiple testing problems that compares expectations of multivariate data among several groups. As a test statistic, a conventional (approximate) $t$-statistic is considered, and we determine its rejection…

统计方法学 · 统计学 2016-12-20 Yoshiyuki Ninomiya , Satoshi Kuriki , Toshihiko Shiroishi , Toyoyuki Takada

The intra-cluster correlation coefficient (ICC) plays an important role while designing the cluster randomized trials (CRTs). Often optimal CRTs are designed assuming that the magnitude of the ICC is constant across the clusters. However,…

统计计算 · 统计学 2019-02-21 Satya Prakash Singh , Pradeep Yadav

Decision trees with binary splits are popularly constructed using Classification and Regression Trees (CART) methodology. For binary classification and regression models, this approach recursively divides the data into two near-homogenous…

机器学习 · 统计学 2020-08-17 Jason M. Klusowski

A new clustering accuracy measure is proposed to determine the unknown number of clusters and to assess the quality of clustering of a data set given in any dimensional space. Our validity index applies the classical nonparametric…

统计方法学 · 统计学 2022-02-15 Soumita Modak

Datasets may contain observations with multiple labels. If the labels are not mutually exclusive, and if the labels vary greatly in frequency, obtaining a sample that includes sufficient observations with scarcer labels to make inferences…

机器学习 · 计算机科学 2026-05-27 Simon Chung , Colby J. Vorland , Donna L. Maney , Andrew W. Brown

Well-spread samples are desirable in many disciplines because they improve estimation when target variables exhibit spatial structure. This paper introduces an integrated methodological framework for spreading samples over the population's…

统计方法学 · 统计学 2025-10-29 Bardia Panahbehagh , Mehdi Mohebbi , Amir Mohammad HosseiniNasab

The Gini's mean difference was defined as the expected absolute difference between a random variable and its independent copy. The corresponding normalized version, namely Gini's index, denotes two times the area between the egalitarian…

概率论 · 数学 2024-01-08 Marco Capaldo , Jorge Navarro

This paper proposes an adaptive random experiment design (ARED) algorithm that can be applied to optimize the multiple factors and levels experiments. The algorithm takes real-time model error as the adaptive condition, and outputs a model…

信号处理 · 电气工程与系统科学 2020-09-01 Zhou Qiao , Duan Xiaochang , Tang Wei

Global survey research increasingly informs high-stakes decisions in AI governance and cross-cultural policy, yet no standardized metric quantifies how well a sample's demographic composition matches its target population. Response rates…

统计方法学 · 统计学 2026-02-23 Evan Hadfield , Andrew Konya

Cross-level interactions among fixed effects in linear mixed models (also known as multilevel models) are often complicated by the variances stemming from random effects and residuals. When these variances change across clusters, tests of…

统计方法学 · 统计学 2022-03-18 Ting Wang , Edgar C. Merkle , Joaquin A. Anguera , Brandon M. Turner

Rerandomization is a modern experimental design technique that repeatedly randomizes treatment assignments until covariates are deemed balanced between treatment groups. This enhances the precision and coherence of causal effect estimators,…

统计方法学 · 统计学 2025-12-08 Antônio Carlos Herling Ribeiro Junior , Zach Branson

Machine learning-driven rankings, where individuals (or items) are ranked in response to a query, mediate search exposure or attention in a variety of safety-critical settings. Thus, it is important to ensure that such rankings are fair.…

机器学习 · 计算机科学 2025-02-18 Aparna Balagopalan , Kai Wang , Olawale Salaudeen , Asia Biega , Marzyeh Ghassemi

Semisupervised methods inevitably invoke some assumption that links the marginal distribution of the features to the regression function of the label. Most commonly, the cluster or manifold assumptions are used which imply that the…

统计理论 · 数学 2011-12-02 Martin Azizyan , Aarti Singh , Larry Wasserman

The Boltzmann-Shannon Index (BSI) for clustered continuous data is introduced as a normalized measure that captures the relationship between geometry-based and frequency-based probability distributions defined over the clusters. In essence,…

信息论 · 计算机科学 2025-12-18 Emanuele Bossi , C. Tyler Diggans , Abd AlRahman R. AlMomani

Identifying relationships between molecular variations and their clinical presentations has been challenged by the heterogeneous causes of a disease. It is imperative to unveil the relationship between the high dimensional molecular…

统计方法学 · 统计学 2021-09-02 Wennan Chang , Changlin Wan , Yong Zang , Chi Zhang , Sha Cao

The performance (accuracy and robustness) of several clustering algorithms is studied for linearly dependent random variables in the presence of noise. It turns out that the error percentage quickly increases when the number of observations…

应用统计 · 统计学 2009-11-13 Pamela Minicozzi , Fabio Rapallo , Enrico Scalas , Francesco Dondero

We present an optimized rerandomization design procedure for a non-sequential treatment-control experiment. Randomized experiments are the gold standard for finding causal effects in nature. But sometimes random assignments result in…

统计方法学 · 统计学 2021-01-26 Adam Kapelner , Abba M. Krieger , Michael Sklar , David Azriel

We introduce a random recursive tree model with two communities, called balanced community modulated random recursive tree, or BCMRT in short. In this setting, pairs of nodes of different type appear sequentially. Each node of the pair…

统计理论 · 数学 2024-02-12 Anna Ben-Hamou , Vasiliki Velona

Determining whether an algorithmic decision-making system discriminates against a specific demographic typically involves comparing a single point estimate of a fairness metric against a predefined threshold. This practice is statistically…

机器学习 · 计算机科学 2026-03-20 Antonio Ferrara , Francesco Cozzi , Alan Perotti , André Panisson , Francesco Bonchi

Model merging has shown that multitask models can be created by directly combining the parameters of different models that are each specialized on tasks of interest. However, models trained independently on distinct tasks often exhibit…

机器学习 · 计算机科学 2026-03-17 Pratik Ramesh , George Stoica , Arun Iyer , Leshem Choshen , Judy Hoffman
‹ 上一页 1 8 9 10 下一页 ›