中文
相关论文

相关论文: Jaccard/Tanimoto similarity test and estimation me…

200 篇论文

Similarity index is an important scientific tool frequently used to determine whether different pairs of entities are similar with respect to some prefixed characteristics. Some standard measures of similarity index include Jaccard index,…

统计方法学 · 统计学 2023-12-19 Srijan Chattopadhyay , Swapnaneel Bhattacharyya

The Jaccard similarity index has often been employed in science and technology as a means to quantify the similarity between two sets. When modified to operate on real-valued values, the Jaccard similarity index can be applied to compare…

数据分析、统计与概率 · 物理学 2024-10-23 Gonzalo Travieso , Alexandre Benatti , Luciano da F. Costa

This paper addresses the problem of estimating the containment and similarity between two sets using only random samples from each set, without relying on sketches of full sets. The study introduces a binomial model for predicting the…

统计计算 · 统计学 2025-07-22 Pranav Joshi

Presence-absence data is defined by vectors or matrices of zeroes and ones, where the ones usually indicate a "presence" in a certain place. Presence-absence data occur for example when investigating geographical species distributions,…

统计方法学 · 统计学 2021-11-24 Gabriele d'Angella , Christian Hennig

Clinical trial simulation (CTS) is critical in new drug development, providing insight into safety and efficacy while guiding trial design. Achieving realistic outcomes in CTS requires an accurately estimated joint distribution of the…

统计方法学 · 统计学 2025-05-08 Longwen Shang , Min Tsao , Xuekui Zhang

The delimitation of biological species, i.e., deciding which individuals belong to the same species and whether and how many different species are represented in a data set, is key to the conservation of biodiversity. Much existing work…

种群与进化 · 定量生物学 2025-12-15 Gabriele d'Angella , Christian Hennig

Binary classification is a task that involves the classification of data into one of two distinct classes. It is widely utilized in various fields. However, conventional classifiers tend to make overconfident predictions for data that…

机器学习 · 计算机科学 2025-03-13 Shoma Yokura , Akihisa Ichiki

Missing covariates are not uncommon in capture-recapture studies. When covariate information is missing at random in capture-recapture data, an empirical full likelihood method has been demonstrated to outperform…

统计方法学 · 统计学 2025-07-15 Yang Liu , Yukun Liu , Pengfei Li , Riquan Zhang

Qualitative interactions occur when a treatment effect or measure of association varies in sign by sub-population. Of particular interest in many biomedical settings are absence/presence qualitative interactions, which occur when an effect…

统计方法学 · 统计学 2020-10-20 Aaron Hudson , Ali Shojaie

We introduce a general class of autoregressive models for studying the dynamic of multivariate binary time series with stationary exogenous covariates. Using a high-level set of assumptions, we show that existence of a stationary path for…

统计理论 · 数学 2024-07-16 Guillaume Franchi , Lionel Truquet

Frequently, empirical studies are plagued with missing data. When the data are missing not at random, the parameter of interest is not identifiable in general. Without additional assumptions, we can derive bounds of the parameters of…

统计方法学 · 统计学 2018-09-12 Zhichao Jiang , Peng Ding

Clinical end-point traits are often characterized by quantitative or qualitative precursors and it has been argued that it may be statistically a more powerful strategy to analyze these precursor traits to decipher the genetic architecture…

统计方法学 · 统计学 2025-04-17 Soumya Sahu , Saurabh Ghosh

To improve confounder adjustments, observational studies are often matched on potential confounders. While matched case-control studies are common and well covered in the literature, our focus here is on matched cohort studies, which are…

Data analyses typically rely upon assumptions about missingness mechanisms that lead to observed versus missing data. When the data are missing not at random, direct assumptions about the missingness mechanism, and indirect assumptions…

统计方法学 · 统计学 2016-03-22 Alexander M Franks , Edoardo M Airoldi , Donald B Rubin

Abundance data are used in ecology for species monitoring and conservation. These count data often display several specific characteristics like numerous missing data, high variance, and a high proportion of zeros, particularly when…

Jaccard Similarity is a very common proximity measurement used to compute the similarity between two asymmetric binary vectors. Jaccard Similarity is the ratio between the 1s (Intersection of two vectors) to 1s (Union of two vectors). This…

数据结构与算法 · 计算机科学 2024-08-20 Varun Puram , Ruthvik Rao Bobbili , Johnson P Thomas

We investigate saddlepoint approximations applied to the score test statistic in genome-wide association studies with binary phenotypes. The inaccuracy in the normal approximation of the score test statistic increases with increasing sample…

统计方法学 · 统计学 2021-10-11 Pål Vegard Johnsen , Øyvind Bakke , Thea Bjørnland , Andrew Thomas DeWan , Mette Langaas

Integrative analysis of datasets generated by multiple cohorts is a widely-used approach for increasing sample size, precision of population estimators, and generalizability of analysis results in epidemiological studies. However, often…

Joint species distribution models are popular in ecology for modeling covariate effects on species occurrence, while characterizing cross-species dependence. Data consist of multivariate binary indicators of the occurrences of different…

统计方法学 · 统计学 2025-07-08 Federica Stolf , David B. Dunson

Collection of genotype data in case-control genetic association studies may often be incomplete for reasons related to genes themselves. This non-ignorable missingness structure, if not appropriately accounted for, can result in…

统计方法学 · 统计学 2024-07-12 Le Wang , Zhengbang Li , Ben Fitzpatrick , Clarice Weinberg , Jinbo Chen
‹ 上一页 1 2 3 10 下一页 ›