English
Related papers

Related papers: Jaccard/Tanimoto similarity test and estimation me…

200 papers

A common problem in clinical trials is to test whether the effect of an explanatory variable on a response of interest is similar between two groups, e.g. patient or treatment groups. In this regard, similarity is defined as equivalence up…

Methodology · Statistics 2024-01-12 Niklas Hagemann , Giampiero Marra , Frank Bretz , Kathrin Möllenhoff

We investigate one/two-sample mean tests for high-dimensional compositional data when the number of variables is comparable with the sample size, as commonly encountered in microbiome research. Existing methods mainly focus on max-type test…

Statistics Theory · Mathematics 2024-04-15 Qianqian Jiang , Wenbo Li , Zeng Li

The Jaccard index is an important similarity measure for item sets and Boolean data. On large datasets, an exact similarity computation is often infeasible for all item pairs both due to time and space constraints, giving rise to faster…

Data Structures and Algorithms · Computer Science 2021-03-09 Marc Bury , Chris Schwiegelshohn , Mara Sorella

So-called 'citizen science' data elicited from crowds has become increasingly popular in many fields including ecology. However, the quality of this information is being frequently debated by many within the scientific community. Therefore,…

Applications · Statistics 2020-05-27 Edgar Santos-Fernandez , Kerrie Mengersen

Statistical matching is a technique for integrating two or more data sets when information available for matching records for individual participants across data sets is incomplete. Statistical matching can be viewed as a missing data…

Methodology · Statistics 2015-10-14 Jae-kwang Kim , Emily Berg , Taesung Park

As the availability of omics data has increased in the last few years, more multi-omics data have been generated, that is, high-dimensional molecular data consisting of several types such as genomic, transcriptomic, or proteomic data, all…

Genomics · Quantitative Biology 2023-02-09 Roman Hornung , Frederik Ludwigs , Jonas Hagenberg , Anne-Laure Boulesteix

Quantifying the similarity of two or more datasets has widespread applications in statistics and machine learning. The method choice is, however, difficult due to the abundance of proposed methods and the lack of neutral comparison studies,…

Methodology · Statistics 2026-04-14 Marieke Stolte , Jörg Rahnenführer , Andrea Bommert

Hybrid controlled trials (HCTs), which augment randomized controlled trials (RCTs) with external controls (ECs), are increasingly receiving attention as a way to address limited power, slow accrual, and ethical concerns in clinical…

Methodology · Statistics 2025-05-02 Jiajun Liu , Ke Zhu , Shu Yang , Xiaofei Wang

The goal of two-sample tests is to assess whether two samples, $S_P \sim P^n$ and $S_Q \sim Q^m$, are drawn from the same distribution. Perhaps intriguingly, one relatively unexplored method to build two-sample tests is the use of binary…

Machine Learning · Statistics 2018-03-14 David Lopez-Paz , Maxime Oquab

Microorganisms are found in almost every environment, including the soil, water, air, and inside other organisms, like animals and plants. While some microorganisms cause diseases, most of them help in biological processes such as…

Machine Learning · Computer Science 2023-09-28 Daniel Agyapong , Jeffrey Ryan Propster , Jane Marks , Toby Dylan Hocking

Presence-only data, point locations where a species has been recorded as being present, are often used in modeling the distribution of a species as a function of a set of explanatory variables---whether to map species occurrence, to…

Applications · Statistics 2010-11-16 David I. Warton , Leah C. Shepherd

With the internet, a massive amount of information on species abundance can be collected under citizen science programs. However, these data are often difficult to use directly in statistical inference, as their collection is generally…

Applications · Statistics 2015-02-27 Christophe Giraud , Clément Calenge , Camille Coron , Romain Julliard

In time-to-event analyses in social sciences, there often exist endogenous time-varying variables, where the event status is correlated with the trajectory of the covariate itself. Ignoring this endogeneity will result in biased estimates.…

Applications · Statistics 2025-04-28 Sophie Potts , Anja Rappl , Karin Kurz , Elisabeth Bergherr

High-dimensional data can be useful for causal inference by providing many confounders that may bolster the plausibility of the ignorability assumption. Propensity score methods are powerful tools for causal inference, are popular in health…

Methodology · Statistics 2017-10-10 Jacob Spertus , Sharon-Lise Normand

Simulation studies are commonly used in methodological research for the empirical evaluation of data analysis methods. They generate artificial data sets under specified mechanisms and compare the performance of methods across conditions.…

Methodology · Statistics 2025-07-11 Samuel Pawel , František Bartoš , Björn S. Siepe , Anna Lohmann

Association testing aims to discover the underlying relationship between genotypes (usually Single Nucleotide Polymorphisms, or SNPs) and phenotypes (attributes, or traits). The typically large data sets used in association testing often…

Applications · Statistics 2012-07-04 Zhen Li , Vikneswaran Gopal , Xiaobo Li , John M. Davis , George Casella

Weakly supervised learning has drawn considerable attention recently to reduce the expensive time and labor consumption of labeling massive data. In this paper, we investigate a novel weakly supervised learning problem of learning from…

Machine Learning · Statistics 2021-02-16 Yuzhou Cao , Lei Feng , Yitian Xu , Bo An , Gang Niu , Masashi Sugiyama

A collaborative distributed binary decision problem is considered. Two statisticians are required to declare the correct probability measure of two jointly distributed memoryless process, denoted by $X^n=(X_1,\dots,X_n)$ and…

Information Theory · Computer Science 2016-04-11 Gil Katz , Pablo Piantanida , Merouane Debbah

Random binnings generated via recursive binary splits are introduced as a way to detect, measure the strength of, and to display the pattern of association between any two variates, whether one or both are continuous or categorical. This…

Methodology · Statistics 2025-04-30 Chris Salahub , Wayne Oldford

During the past few decades, missing-data problems have been studied extensively, with a focus on the ignorable missing case, where the missing probability depends only on observable quantities. By contrast, research into non-ignorable…

Methodology · Statistics 2019-08-06 Yukun Liu , Pengfei Li , Jing Qin