中文
相关论文

相关论文: Jaccard/Tanimoto similarity test and estimation me…

200 篇论文

Data collected in clinical trials are often composed of multiple types of variables. For example, laboratory measurements and vital signs are longitudinal data of continuous or categorical variables, adverse events may be recurrent events,…

统计方法学 · 统计学 2023-01-12 Tuo Wang , Rachel Zilinskas , Ying Li , Yongming Qu

In data mining, when binary prediction rules are used to predict a binary outcome, many performance measures are used in a vast array of literature for the purposes of evaluation and comparison. Some examples include classification…

机器学习 · 统计学 2025-07-08 Zheng Yuan , Wenxin Jiang

We propose a data-driven method to learn the time-dependent probability density of a multivariate stochastic process from sample paths, assuming that the initial probability density is known and can be evaluated. Our method uses a novel…

机器学习 · 统计学 2025-06-19 Agnimitra Dasgupta , Javier Murgoitio-Esandi , Ali Fardisi , Assad A Oberai

Accurate power and sample size estimation are crucial to the design and analysis of genetic association studies. When analyzing a binary trait via logistic regression, important covariates such as age and sex are typically included in the…

统计方法学 · 统计学 2022-10-05 Ziang Zhang , Lei Sun

The probability Jaccard similarity was recently proposed as a natural generalization of the Jaccard similarity to measure the proximity of sets whose elements are associated with relative frequencies or probabilities. In combination with a…

数据结构与算法 · 计算机科学 2020-10-27 Otmar Ertl

We introduce a sign-aware, multistate Jaccard/Tanimoto framework that extends overlap-based distances from nonnegative vectors and measures to arbitrary real- and complex-valued signals while retaining bounded metric and…

机器学习 · 计算机科学 2025-12-24 Vineet Yadav

In epidemiological cohort studies, the relative risk (also known as risk ratio) is a major measure of association to summarize the results of two treatments or exposures. Generally, it measures the relative change in disease risk as a…

统计方法学 · 统计学 2022-07-05 Gopal Nath , Krishna K. Saha , Suojin Wang

Binary classification is a fundamental task in machine learning, with applications spanning various scientific domains. Whether scientists are conducting fundamental research or refining practical applications, they typically assess and…

机器学习 · 计算机科学 2023-10-20 Attila Fazekas , György Kovács

Large-scale population-level datasets, such as the UK Biobank and the All of Us Research Program, often lack covariates needed for a specific analysis, such as genetic or lifestyle measures, while related studies measure them. This creates…

统计方法学 · 统计学 2026-05-07 Huali Zhao , Tianying Wang

Models for accurately predicting species distributions have become essential tools for many ecological and conservation problems. For many species, presence-background (presence-only) data is the most commonly available type of spatial…

统计方法学 · 统计学 2018-01-08 Yan Wang , Lewi Stone

Understanding and predicting interactions between predators and prey and their environment are fundamental for understanding food web structure, dynamics, and ecosystem function in both terrestrial and marine ecosystems.Thus, estimating the…

统计方法学 · 统计学 2023-08-29 H. Solvang , S. Imori , M. Biuw , U. Lindstrøm , T. Haug

Site occupancy models are routinely used to estimate the probability of species presence from either abundance or presence-absence data collected across sites with repeated sampling occasions. In the last two decades, a broad class of…

统计方法学 · 统计学 2022-04-05 Wen-Han Hwang , Jakub Stoklosa , Lu-Fang Chen

Conducting valid statistical analyses is challenging in the presence of missing-not-at-random (MNAR) data, where the missingness mechanism is dependent on the missing values themselves even conditioned on the observed data. Here, we…

统计方法学 · 统计学 2023-06-13 Anna Guo , Jiwei Zhao , Razieh Nabi

Missing data can lead to inefficiencies and biases in analyses, in particular when data are missing not at random (MNAR). It is thus vital to understand and correctly identify the missing data mechanism. Recovering missing values through a…

统计方法学 · 统计学 2022-12-08 Jack Noonan , Adetola Adedamola Adediran , Robin Mitra , Stefanie Biedermann

Count outcomes in longitudinal studies are frequent in clinical and engineering studies. In frequentist and Bayesian statistical analysis, methods such as Mixed linear models allow the variability or correlation within individuals to be…

统计方法学 · 统计学 2024-07-15 Alejandra Estefanía Patiño Hoyos , Johnatan Cardona Jiménez

Survival analysis aims to explore the relationship between covariates and the time until the occurrence of an event. The Cox proportional hazards model is commonly used for right-censored data, but it is not strictly limited to this type of…

统计方法学 · 统计学 2025-07-02 Abdoulaye Dioni , Lynne Moore , Aida Eslami

Testing effect size homogeneity is an essential part when conducting a meta-analysis. Comparative studies of effect size homogeneity tests in case of binary outcomes are found in the literature, but no test has come out as an absolute…

统计方法学 · 统计学 2022-03-11 Osama Almalik

We study an EM algorithm for estimating product-term regression models with missing data. The study of such problems in the likelihood tradition has thus far been restricted to an EM algorithm method using full numerical integration.…

统计方法学 · 统计学 2021-11-16 Dale S. Kim

In many real-world scenarios, interested variables are often represented as discretized values due to measurement limitations. Applying Conditional Independence (CI) tests directly to such discretized data, however, can lead to incorrect…

人工智能 · 计算机科学 2025-06-11 Boyang Sun , Yu Yao , Xinshuai Dong , Zongfang Liu , Tongliang Liu , Yumou Qiu , Kun Zhang

In the current era of systems biological research there is a need for the integrative analysis of binary and quantitative genomics data sets measured on the same objects. One standard tool of exploring the underlying dependence structure…