English
Related papers

Related papers: A Multivariate Equivalence Test Based on Mahalanob…

200 papers

The past decade has witnessed the rapid development of feature representation learning and distance metric learning, whereas the two steps are often discussed separately. To explore their interaction, this work proposes an end-to-end…

Computer Vision and Pattern Recognition · Computer Science 2016-04-18 Guangrun Wang , Liang Lin , Shengyong Ding , Ya Li , Qing Wang

Distance metric learning is of fundamental interest in machine learning because the distance metric employed can significantly affect the performance of many learning methods. Quadratic Mahalanobis metric learning is a popular approach to…

Machine Learning · Computer Science 2013-02-15 Chunhua Shen , Junae Kim , Fayao Liu , Lei Wang , Anton van den Hengel

The need for appropriate ways to measure the distance or similarity between data is ubiquitous in machine learning, pattern recognition and data mining, but handcrafting such good metrics for specific problems is generally difficult. This…

Machine Learning · Computer Science 2019-01-25 Aurélien Bellet , Amaury Habrard , Marc Sebban

Data depth has been applied as a nonparametric measurement for ranking multivariate samples. In this paper, we focus on homogeneity tests to assess whether two multivariate samples are from the same distribution. There are many data…

Statistics Theory · Mathematics 2023-06-09 Yiting Chen , Wei Lin , Xiaoping Shi

Clinical trials often aim to compare a new drug with a reference treatment in terms of efficacy and/or toxicity depending on covariates such as, for example, the dose level of the drug. Equivalence of these treatments can be claimed if the…

Methodology · Statistics 2019-10-22 Holger Dette , Kathrin Möllenhoff , Frank Bretz

We study optimal sample allocation between treatment and control groups under Bayesian linear models. We derive an analytic expression for the Bayes risk, which depends jointly on sample size and covariate mean balance across groups. Under…

Statistics Theory · Mathematics 2025-09-03 André A. F. Fumis , Victor Fossaluza , Rafael B. Stern

This report provides an exploration of different distance measures that can be used with the $K$-means algorithm for cluster analysis. Specifically, we investigate the Mahalanobis distance, and critically assess any benefits it may have…

Other Statistics · Statistics 2024-04-23 Zoe Shapcott

Two-sample hypothesis testing-determining whether two sets of data are drawn from the same distribution-is a fundamental problem in statistics and machine learning with broad scientific applications. In the context of nonparametric testing,…

Machine Learning · Statistics 2026-04-21 Antoine Chatalic , Marco Letizia , Nicolas Schreuder , Lorenzo Rosasco

How well can multiple incompatible observables be implemented by a single measurement? This is a fundamental problem in quantum mechanics with wide implications for the performance optimization of numerous tasks in quantum information…

Quantum Physics · Physics 2024-10-10 Hongzhen Chen , Lingna Wang , Haidong Yuan

Finite mixtures of multivariate normal distributions have been widely used in empirical applications in diverse fields such as statistical genetics and statistical finance. Testing the number of components in multivariate normal mixture…

Statistics Theory · Mathematics 2019-02-11 Hiroyuki Kasahara , Katsumi Shimotsu

The aim of this thesis is to find a solution to the non-parametric independence problem in separable metric spaces. Suppose we are given finite collection of samples from an i.i.d. sequence of paired random elements, where each marginal has…

Statistics Theory · Mathematics 2017-06-13 Martin Emil Jakobsen

Log-Euclidean distances are commonly used to quantify the similarity between positive definite matrices using geometric considerations. This paper analyzes the behavior of this distance when it is used to measure closeness between…

Signal Processing · Electrical Eng. & Systems 2024-08-09 Xavier Mestre , Roberto Pereira

There is a growing interest in societal concerns in machine learning systems, especially in fairness. Multicalibration gives a comprehensive methodology to address group fairness. In this work, we address the multicalibration error and…

Machine Learning · Computer Science 2021-06-08 Eliran Shabat , Lee Cohen , Yishay Mansour

Labelled Markov chains (LMCs) are widely used in probabilistic verification, speech recognition, computational biology, and many other fields. Checking two LMCs for equivalence is a classical problem subject to extensive studies, while the…

Logic in Computer Science · Computer Science 2014-05-16 Taolue Chen , Stefan Kiefer

The distance standard deviation, which arises in distance correlation analysis of multivariate data, is studied as a measure of spread. The asymptotic distribution of the empirical distance standard deviation is derived under the assumption…

Statistics Theory · Mathematics 2019-12-12 Dominic Edelmann , Donald Richards , Daniel Vogel

Due to rapid technological advances, a wide range of different measurements can be obtained from a given biological sample including single nucleotide polymorphisms, copy number variation, gene expression levels, DNA methylation and…

Methodology · Statistics 2013-03-29 Christopher Minas , Edward Curry , Giovanni Montana

Data collected in clinical trials are often composed of multiple types of variables. For example, laboratory measurements and vital signs are longitudinal data of continuous or categorical variables, adverse events may be recurrent events,…

Methodology · Statistics 2023-01-12 Tuo Wang , Rachel Zilinskas , Ying Li , Yongming Qu

In several modern applications, ranging from genetics to genomics and neuroimaging, there is a need to compare observations across different populations, such as groups of healthy and diseased individuals. The interest is in detecting a…

Methodology · Statistics 2012-05-14 Christopher Minas , Giovanni Montana

Protecting individual privacy is essential across research domains, from socio-economic surveys to big-tech user data. This need is particularly acute in healthcare, where analyses often involve sensitive patient information. A typical…

Applications · Statistics 2026-04-09 Savita Pareek , Luca Insolia , Roberto Molinari , Stéphane Guerrier

The seminal work of Morgan and Rubin (2012) considers rerandomization for all the units at one time. In practice, however, experimenters may have to rerandomize units sequentially. For example, a clinician studying a rare disease may be…

Applications · Statistics 2018-04-17 Quan Zhou , Philip Ernst , Kari Lock Morgan , Donald Rubin , Anru Zhang