English
Related papers

Related papers: A New Burrows Wheeler Transform Markov Distance

200 papers

Time series are high-dimensional and complex data objects, making their efficient search and indexing a longstanding challenge in data mining. Building on a recently introduced similarity measure, namely Multiscale Dubuc Distance (MDD),…

Machine Learning · Computer Science 2025-10-28 Azim Ahmadzadeh , Mahsa Khazaei , Elaina Rohlfing

The literature postulates that the dynamic time warping (dtw) distance can cope with temporal variations but stores and processes time series in a form as if the dtw-distance cannot cope with such variations. To address this inconsistency,…

Machine Learning · Computer Science 2019-03-11 Brijnesh Jain

Motivated by the challenge of using DNA-seq data to identify viruses in human blood samples, we propose a novel classification algorithm called "Radial Distance Weighted Discrimination" (or Radial DWD). This classifier is designed for…

Applications · Statistics 2016-02-10 Jie Xiong , D. P. Dittmer , J. S. Marron

The boom of genomic sequencing makes compression of set of sequences inescapable. This underlies the need for multi-string indexing data structures that helps compressing the data. The most prominent example of such data structures is the…

Data Structures and Algorithms · Computer Science 2021-11-18 Bastien Cazaux , Eric Rivals

High dimension low sample size statistical analysis is important in a wide range of applications. In such situations, the highly appealing discrimination method, support vector machine, can be improved to alleviate data piling at the…

Optimization and Control · Mathematics 2017-08-18 Xin Yee Lam , J. S. Marron , Defeng Sun , Kim-Chuan Toh

Normalized Compression Distance (NCD) is a popular tool that uses compression algorithms to cluster and classify data in a wide range of applications. Existing discussions of NCD's theoretical merit rely on certain theoretical properties of…

Cryptography and Security · Computer Science 2015-09-03 Rebecca Schuller Borbely

Current advances in next generation sequencing techniques have allowed researchers to conduct comprehensive research on microbiome and human diseases, with recent studies identifying associations between human microbiome and health outcomes…

Methodology · Statistics 2021-06-09 Konstantin Shestopaloff , Mei Dong , Fan Gao , Wei Xu

We present a new method for clustering based on compression. The method doesn't use subject-specific features or background knowledge, and works as follows: First, we determine a universal similarity distance, the normalized compression…

Computer Vision and Pattern Recognition · Computer Science 2007-05-23 Rudi Cilibrasi , Paul Vitanyi

Gaussian mixture models (GMMs) are widely used in machine learning for tasks such as clustering, classification, image reconstruction, and generative modeling. A key challenge in working with GMMs is defining a computationally efficient and…

Machine Learning · Computer Science 2025-08-05 Moritz Piening , Robert Beinert

Collaborative filtering, a widely-used recommendation technique, predicts a user's preference by aggregating the ratings from similar users. As a result, these measures cannot fully utilize the rating information and are not suitable for…

Information Retrieval · Computer Science 2019-12-11 Yitong Meng , Xinyan Dai , Xiao Yan , James Cheng , Weiwen Liu , Benben Liao , Jun Guo , Guangyong Chen

Measuring distance or similarity between time-series data is a fundamental aspect of many applications including classification, clustering, and ensembling/alignment. Existing measures may fail to capture similarities among local trends…

Machine Learning · Computer Science 2024-12-20 Ajitesh Srivastava

We consider the paradigm of unsupervised anomaly detection, which involves the identification of anomalies within a dataset in the absence of labeled examples. Though distance-based methods are top-performing for unsupervised anomaly…

Machine Learning · Statistics 2025-08-04 Yuchao Cai , Hanfang Yang , Yuheng Ma , Hanyuan Hang

The word mover's distance (WMD) is a popular semantic similarity metric for two texts. This position paper studies several possible extensions of WMD. We experiment with the frequency of words in the corpus as a weighting factor and the…

Computation and Language · Computer Science 2022-02-09 Ilya Smirnov , Ivan P. Yamshchikov

Measuring similarities between unlabeled time series trajectories is an important problem in domains as diverse as medicine, astronomy, finance, and computer vision. It is often unclear what is the appropriate metric to use because of the…

Machine Learning · Computer Science 2018-10-25 Abubakar Abid , James Zou

Manifold distances are very effective tools for visual object recognition. However, most of the traditional manifold distances between images are based on the pixel-level comparison and thus easily affected by image rotations and…

Computer Vision and Pattern Recognition · Computer Science 2016-05-13 Fengfu Li , Xiayuan Huang , Hong Qiao , Bo Zhang

The proliferation of malware variants poses a significant challenges to traditional malware detection approaches, such as signature-based methods, necessitating the development of advanced machine learning techniques. In this research, we…

Machine Learning · Computer Science 2024-12-30 Ritik Mehta , Olha Jureckova , Mark Stamp

In this paper we describe algorithms for computing the BWT and for building (compressed) indexes in external memory. The innovative feature of our algorithms is that they are lightweight in the sense that, for an input of size $n$, they use…

Data Structures and Algorithms · Computer Science 2009-09-25 Paolo Ferragina , Travis Gagie , Giovanni Manzini

The maximum mean discrepancy and Wasserstein distance are popular distance measures between distributions and play important roles in many machine learning problems such as metric learning, generative modeling, domain adaption, and…

Machine Learning · Computer Science 2025-01-22 Dong Qiao , Jicong Fan

Modern data often take the form of a multiway array. However, most classification methods are designed for vectors, i.e., 1-way arrays. Distance weighted discrimination (DWD) is a popular high-dimensional classification method that has been…

Methodology · Statistics 2021-10-12 Bin Guo , Lynn E. Eberly , Pierre-Gilles Henry , Christophe Lenglet , Eric F. Lock

This paper is concerned with the problem of recovering third-order tensor data from limited samples. A recently proposed tensor decomposition (BMD) method has been shown to efficiently compress third-order spatiotemporal data. Using the…

Numerical Analysis · Mathematics 2024-02-21 Fan Tian , Mirjeta Pasha , Misha E. Kilmer , Eric Miller , Abani Patra