English
Related papers

Related papers: A robust, scalable K-statistic for quantifying imm…

200 papers

On the smallest scales, three-dimensional large-scale structure surveys contain a wealth of cosmological information which cannot be trivially extracted due to the non-linear dynamical evolution of the density field. Lagrangian perturbation…

Cosmology and Nongalactic Astrophysics · Physics 2015-08-05 Florent Leclercq , Jens Jasche , Héctor Gil-Marín , Benjamin Wandelt

Partitioning Around Medoids (PAM, k-Medoids) is a popular clustering technique to use with arbitrary distance functions or similarities, where each cluster is represented by its most central object, called the medoid or the discrete median.…

Machine Learning · Computer Science 2023-09-07 Lars Lenssen , Erich Schubert

The intensity function and Ripley's K-function have been used extensively in the literature to describe the first and second moment structure of spatial point sets. This has many applications including describing the statistical structure…

Methodology · Statistics 2018-12-18 Jon Sporring , Rasmus Waagepetersen , Stefan Sommer

In various applications with large spatial regions, the relationship between the response variable and the covariates is expected to exhibit complex spatial patterns. We propose a spatially clustered varying coefficient model, where the…

Methodology · Statistics 2020-07-21 Fangzheng Lin , Yanlin Tang , Huichen Zhu , Zhongyi Zhu

Next location prediction underpins a growing number of mobility, retail, and public-health applications, yet its societal impacts remain largely unexplored. In this paper, we audit state-of-the-art mobility prediction models trained on a…

Machine Learning · Computer Science 2025-11-03 Ashwin Kumar , Hanyu Zhang , David A. Schweidel , William Yeoh

Similarity search based on a distance function in metric spaces is a fundamental problem for many applications. Queries for similar objects lead to the well-known machine learning task of nearest-neighbours identification. Many data…

Information Retrieval · Computer Science 2022-08-05 Felipe Ortega , Maria Jesus Algar , Isaac Martín de Diego , Javier M. Moguerza

Szapudi et al (2001) introduced the method of estimating angular power spectrum of the CMB sky via heuristically weighted correlation functions. Part of the new technique is that all (co)variances are evaluated by massive Monte Carlo…

Astrophysics · Physics 2007-05-23 I. Szapudi , S. Prunet , S. Colombi

The study of immune cellular composition has been of great scientific interest in immunology because of the generation of multiple large-scale data. From the statistical point of view, such immune cellular data should be treated as…

Applications · Statistics 2022-04-22 Jinkyung Yoo , Zequn Sun , Michael Greenacre , Qin Ma , Dongjun Chung , Young Min Kim

Given a point set S and an unknown metric d on S, we study the problem of efficiently partitioning S into k clusters while querying few distances between the points. In our model we assume that we have access to one versus all queries that…

Data Structures and Algorithms · Computer Science 2011-05-10 Konstantin Voevodski , Maria-Florina Balcan , Heiko Roglin , Shang-Hua Teng , Yu Xia

Pragmatic trials evaluating health care interventions often adopt cluster randomization due to scientific or logistical considerations. Previous reviews have shown that co-primary endpoints are common in pragmatic trials but infrequently…

Methodology · Statistics 2022-05-03 Siyun Yang , Mirjam Moerbeek , Monica Taljaard , Fan Li

Spatial transcriptomics measures the expression of thousands of genes in a tissue sample while preserving its spatial structure. This class of technologies has enabled the investigation of the spatial variation of gene expressions and their…

Methodology · Statistics 2025-10-23 Andrea Sottosanti , Davide Risso , Francesco Denti

The $k$-means is one of the most important unsupervised learning techniques in statistics and computer science. The goal is to partition a data set into many clusters, such that observations within clusters are the most homogeneous and…

Machine Learning · Statistics 2022-11-21 Tonglin Zhang

This paper presents a method for robust optimization for online incremental Simultaneous Localization and Mapping (SLAM). Due to the NP-Hardness of data association in the presence of perceptual aliasing, tractable (approximate) approaches…

Robotics · Computer Science 2023-04-28 Daniel McGann , John G. Rogers , Michael Kaess

In this paper, we first propose a new iterative algorithm, called the K-sets+ algorithm for clustering data points in a semi-metric space, where the distance measure does not necessarily satisfy the triangular inequality. We show that the…

Data Structures and Algorithms · Computer Science 2017-05-12 Cheng-Shang Chang , Chia-Tai Chang , Duan-Shin Lee , Li-Heng Liou

Clustering is an important exploratory data analysis technique to group objects based on their similarity. The widely used $K$-means clustering method relies on some notion of distance to partition data into a fewer number of groups. In the…

Machine Learning · Statistics 2022-10-14 Yubo Zhuang , Xiaohui Chen , Yun Yang

Kepler and K2 data analysis reported in the literature is mostly based on aperture photometry. Because of Kepler's large, undersampled pixels and the presence of nearby sources, aperture photometry is not always the ideal way to obtain…

Solar and Stellar Astrophysics · Physics 2016-01-27 M. Libralato , L. R. Bedin , D. Nardiello , G. Piotto

We study sequences of scaled edge-corrected empirical (generalized) K-functions (modifying Ripley's K-function) each of them constructed from a single observation of a $d$-dimensional fourth-order stationary point process in a sampling…

Statistics Theory · Mathematics 2017-06-06 Lothar Heinrich

Not-at-random missingness presents a challenge in addressing missing data in many health research applications. In this paper, we propose a new approach to account for not-at-random missingness after multiple imputation through weighted…

Methodology · Statistics 2021-01-21 Lauren J Beesley , Jeremy M G Taylor

The analysis of area-level aggregated summary data is common in many disciplines including epidemiology and the social sciences. Typically, Markov random field spatial models have been employed to acknowledge spatial dependence and allow…

Methodology · Statistics 2017-09-28 Katherine Wilson , Jon Wakefield

A spatially distributed system contains a large amount of agents with limited sensing, data processing, and communication capabilities. Recent technological advances have opened up possibilities to deploy spatially distributed systems for…

Information Theory · Computer Science 2015-11-30 Cheng Cheng , Yingchun Jiang , Qiyu Sun