English
Related papers

Related papers: Jaccard/Tanimoto similarity test and estimation me…

200 papers

Motivated by investigating spatio-temporal patterns of the distribution of continuous variables, we consider describing the conditional distribution function of the response variable incorporating spatio-temporal components given…

The ICH E9 addendum introduces the term intercurrent event to refer to events that happen after randomisation and that can either preclude observation of the outcome of interest or affect its interpretation. It proposes five strategies for…

Methodology · Statistics 2021-07-12 Camila Olarte Parra , Rhian M. Daniel , Jonathan W. Bartlett

Statistical models for describing the probability distribution over the states of biological systems are commonly used for dimensional reduction. Among these models, pairwise models are very attractive in part because they can be fit using…

Quantitative Methods · Quantitative Biology 2009-11-30 Yasser Roudi , Erik Aurell , John Hertz

Blockwise missing data occurs frequently when we integrate multisource or multimodality data where different sources or modalities contain complementary information. In this paper, we consider a high-dimensional linear regression model with…

Methodology · Statistics 2023-06-30 Fei Xue , Rong Ma , Hongzhe Li

Imputing missing values is an important preprocessing step in data analysis, but the literature offers little guidance on how to choose between different imputation models. This letter suggests adopting the imputation model that generates a…

Methodology · Statistics 2021-07-13 Moritz Marbach

We study the identification and estimation of statistical functionals of multivariate data missing non-monotonically and not-at-random, taking a semiparametric approach. Specifically, we assume that the missingness mechanism satisfies what…

Methodology · Statistics 2022-12-26 Daniel Malinsky , Ilya Shpitser , Eric J Tchetgen Tchetgen

We compare the performance of standard nearest-neighbor propensity score matching with that of an analogous Bayesian propensity score matching procedure. We show that the Bayesian approach makes better use of available information, as it…

Methodology · Statistics 2021-05-07 R. Michael Alvarez , Ines Levin

Identifying which taxa in our microbiota are associated with traits of interest is important for advancing science and health. However, the identification is challenging because the measured vector of taxa counts (by amplicon sequencing) is…

Genomics · Quantitative Biology 2020-03-31 Barak Brill , Amnon Amir , Ruth Heller

Accurate and precise covariance matrices will be important in enabling planned cosmological surveys to detect new physics. Standard methods imply either the need for many N-body simulations in order to obtain an accurate estimate, or a…

Cosmology and Nongalactic Astrophysics · Physics 2018-12-13 Alex Hall , Andy Taylor

Commonly used methods to analyze incomplete longitudinal clinical trial data include complete case analysis (CC) and last observation carried forward (LOCF). However, such methods rest on strong assumptions, including missing completely at…

Statistics Theory · Mathematics 2007-06-13 Ivy Jansen , Caroline Beunckens , Geert Molenberghs , Geert Verbeke , Craig Mallinckrodt

The discovery of inhabited exoplanets hinges on identifying biosignature gases. JWST can reveal biosignature gases, though current discoveries have yet to evidence life. The central challenge is attribution: how can we confidently identify…

Earth and Planetary Astrophysics · Physics 2026-02-17 Tereza Constantinou , Oliver Shorttle , Miles Cranmer , Paul B. Rimmer

We introduce equivalence testing procedures for linear regression analyses. Such tests can be very useful for confirming the lack of a meaningful association between a continuous outcome and a continuous or binary predictor. Specifically,…

Methodology · Statistics 2023-05-17 Harlan Campbell

In epidemiological studies of time-to-event data, a quantity of interest to the clinician and the patient is the risk of an event given a covariate profile. However, methods relying on time matching or risk-set sampling (including Cox…

Methodology · Statistics 2020-09-23 Sahir Rai Bhatnagar , Maxime Turgeon , Jesse Islam , James A. Hanley , Olli Saarela

Many practical studies rely on hypothesis testing procedures applied to data sets with missing information. An important part of the analysis is to determine the impact of the missing data on the performance of the test, and this can be…

Methodology · Statistics 2011-02-15 Dan L. Nicolae , Xiao-Li Meng , Augustine Kong

Missing data arise in most applied settings and are ubiquitous in electronic health records (EHR). When data are missing not at random (MNAR) with respect to measured covariates, sensitivity analyses are often considered. These post-hoc…

Methodology · Statistics 2023-07-11 Alexander W. Levis , Rajarshi Mukherjee , Rui Wang , Heidi Fischer , Sebastien Haneuse

Data similarity (or distance) computation is a fundamental research topic which underpins many high-level applications based on similarity measures in machine learning and data mining. However, in large-scale real-world scenarios, the exact…

Data Structures and Algorithms · Computer Science 2018-11-13 Wei Wu , Bin Li , Ling Chen , Junbin Gao , Chengqi Zhang

Accurately predicting the geographic ranges of species is crucial for assisting conservation efforts. Traditionally, range maps were manually created by experts. However, species distribution models (SDMs) and, more recently, deep…

Quantitative Methods · Quantitative Biology 2024-08-29 Filip Dorm , Christian Lange , Scott Loarie , Oisin Mac Aodha

Complex high-dimensional co-occurrence data are increasingly popular from a complex system of interacting physical, biological and social processes in discretely indexed modifiable areal units or continuously indexed locations of a study…

Machine Learning · Statistics 2022-10-19 Jian-Yong Wang , Han Yu

In large scale genetic association studies, a primary aim is to test for association between genetic variants and a disease outcome. The variants of interest are often rare, and appear with low frequency among subjects. In this situation,…

Methodology · Statistics 2017-12-20 Arjun Sondhi , Kenneth Martin Rice

In this paper, the authors first provide an overview of two major developments on complex survey data analysis: the empirical likelihood methods and statistical inference with non-probability survey samples, and highlight the important…

Methodology · Statistics 2025-08-14 Yilin Chen , Pengfei Li , J. N. K. Rao , Changbao Wu