English
Related papers

Related papers: Compression ratios based on the Universal Similari…

200 papers

The Universal Similarity Metric (USM) has been demonstrated to give practically useful measures of "similarity" between sequence data. Here we have used the USM as an alternative distance metric in a K-Nearest Neighbours (K-NN) learner to…

Machine Learning · Computer Science 2024-05-13 David Lindsay , Sian Lindsay

We present a new method for clustering based on compression. The method doesn't use subject-specific features or background knowledge, and works as follows: First, we determine a universal similarity distance, the normalized compression…

Computer Vision and Pattern Recognition · Computer Science 2007-05-23 Rudi Cilibrasi , Paul Vitanyi

We present a new similarity measure based on information theoretic measures which is superior than Normalized Compression Distance for clustering problems and inherits the useful properties of conditional Kolmogorov complexity. We show that…

Machine Learning · Statistics 2014-10-22 Andrey Bogomolov , Bruno Lepri , Fabio Pianesi

A new class of distances appropriate for measuring similarity relations between sequences, say one type of similarity per distance, is studied. We propose a new ``normalized information distance'', based on the noncomputable notion of…

Computational Complexity · Computer Science 2011-11-09 Ming Li , Xin Chen , Xin Li , Bin Ma , Paul Vitanyi

Enormous volumes of short reads data from next-generation sequencing (NGS) technologies have posed new challenges to the area of genomic sequence comparison. The multiple sequence alignment approach is hardly applicable to NGS data due to…

Genomics · Quantitative Biology 2020-03-25 Ngoc Hieu Tran , Xin Chen

Distance-based clustering and classification are widely used in various fields to group mixed numeric and categorical data. In many algorithms, a predefined distance measurement is used to cluster data points based on their dissimilarity.…

Machine Learning · Computer Science 2024-10-14 Jesse S. Ghashti , John R. J. Thompson

The advent of high dimensional single cell data in the biomedical sciences has necessitated the development of dimensionality-reduction tools. t-SNE and UMAP are the two most frequently used approaches, allowing clear visualisation of…

Quantitative Methods · Quantitative Biology 2021-12-09 Carlos P. Roca1 , Oliver T. Burton , Julika Neumann , Samar Tareen , Carly E. Whyte , Stéphanie Humblet-Baron , Adrian Liston

We study the problem of efficiently clustering protein sequences in a limited information setting. We assume that we do not know the distances between the sequences in advance, and must query them during the execution of the algorithm. Our…

Data Structures and Algorithms · Computer Science 2015-03-17 Konstantin Voevodski , Maria-Florina Balcan , Heiko Roglin , Shang-Hua Teng , Yu Xia

A multitude of measures have been proposed to quantify the similarity between protein 3-D structure. Among these measures, contact map overlap (CMO) maximization deserved sustained attention during past decade because it offers a fine…

Quantitative Methods · Quantitative Biology 2009-04-20 Rumen Andonov , Nicola Yanev , Noël Malod-Dognin

Single-cell omics enable the profiles of cells, which contain large numbers of biological features, to be quantified. Cluster analysis, a dimensionality reduction process, is used to reduce the dimensions of the data to make it…

Genomics · Quantitative Biology 2024-06-06 Okezue Bell , Arthur Lee , Elizabeth Engle

We study the problem of how well a tree metric is able to preserve the sum of pairwise distances of an arbitrary metric. This problem is closely related to low-stretch metric embeddings and is interesting by its own flavor from the line of…

Data Structures and Algorithms · Computer Science 2013-01-16 Mong-Jen Kao , Der-Tsai Lee , Dorothea Wagner

Normalized compression distance (NCD) is a parameter-free, feature-free, alignment-free, similarity measure between a pair of finite objects based on compression. However, it is not sufficient for all applications. We propose an NCD of…

Computer Vision and Pattern Recognition · Computer Science 2016-01-28 Andrew R. Cohen , Paul M. B. Vitanyi

Research into the classification of time series has made enormous progress in the last decade. The UCR time series archive has played a significant role in challenging and guiding the development of new learners for time series…

We consider the problem of metric learning subject to a set of constraints on relative-distance comparisons between the data items. Such constraints are meant to reflect side-information that is not expressed directly in the feature vectors…

Machine Learning · Computer Science 2016-12-06 Ehsan Amid , Aristides Gionis , Antti Ukkonen

In another related work, U-statistics were used for non-asymptotic "average-case" analysis of random compressed sensing matrices. In this companion paper the same analytical tool is adopted differently - here we perform non-asymptotic…

Information Theory · Computer Science 2015-06-11 Fabian Lim , Vladimir Stojanovic

Cosmic shear, galaxy clustering, and the abundance of massive halos each probe the large-scale structure of the Universe in complementary ways. We present cosmological constraints from the joint analysis of the three probes, building on the…

Cosmology and Nongalactic Astrophysics · Physics 2025-03-14 S. Bocquet , S. Grandis , E. Krause , C. To , L. E. Bleem , M. Klein , J. J. Mohr , T. Schrabback , A. Alarcon , O. Alves , A. Amon , F. Andrade-Oliveira , E. J. Baxter , K. Bechtol , M. R. Becker , G. M. Bernstein , J. Blazek , H. Camacho , A. Campos , A. Carnero Rosell , M. Carrasco Kind , R. Cawthon , C. Chang , R. Chen , A. Choi , J. Cordero , M. Crocce , C. Davis , J. DeRose , H. T. Diehl , S. Dodelson , C. Doux , A. Drlica-Wagner , K. Eckert , T. F. Eifler , F. Elsner , J. Elvin-Poole , S. Everett , X. Fang , A. Ferté , P. Fosalba , O. Friedrich , J. Frieman , M. Gatti , G. Giannini , D. Gruen , R. A. Gruendl , I. Harrison , W. G. Hartley , K. Herner , H. Huang , E. M. Huff , D. Huterer , M. Jarvis , N. Kuropatkin , P. -F. Leget , P. Lemos , A. R. Liddle , N. MacCrann , J. McCullough , J. Muir , J. Myles , A. Navarro-Alsina , S. Pandey , Y. Park , A. Porredon , J. Prat , M. Raveri , R. P. Rollins , A. Roodman , R. Rosenfeld , E. S. Rykoff , C. Sánchez , J. Sanchez , L. F. Secco , I. Sevilla-Noarbe , E. Sheldon , T. Shin , M. A. Troxel , I. Tutusaus , T. N. Varga , N. Weaverdyck , R. H. Wechsler , H. -Y. Wu , B. Yanny , B. Yin , Y. Zhang , J. Zuntz , T. M. C. Abbott , P. A. R. Ade , M. Aguena , S. Allam , S. W. Allen , A. J. Anderson , B. Ansarinejad , J. E. Austermann , M. Bayliss , J. A. Beall , A. N. Bender , B. A. Benson , F. Bianchini , M. Brodwin , D. Brooks , L. Bryant , D. L. Burke , R. E. A. Canning , J. E. Carlstrom , J. Carretero , F. J. Castander , C. L. Chang , P. Chaubal , H. C. Chiang , T-L. Chou , R. Citron , C. Corbett Moran , M. Costanzi , T. M. Crawford , A. T. Crites , L. N. da Costa , M. E. S. Pereira , T. M. Davis , T. de Haan , M. A. Dobbs , P. Doel , W. Everett , A. Farahi , B. Flaugher , A. M. Flores , B. Floyd , J. Gallicchio , E. Gaztanaga , E. M. George , M. D. Gladders , N. Gupta , G. Gutierrez , N. W. Halverson , S. R. Hinton , J. Hlavacek-Larrondo , G. P. Holder , D. L. Hollowood , W. L. Holzapfel , J. D. Hrubes , N. Huang , J. Hubmayr , K. D. Irwin , D. J. James , F. Kéruzoré , G. Khullar , K. Kim , L. Knox , R. Kraft , K. Kuehn , O. Lahav , A. T. Lee , S. Lee , D. Li , C. Lidman , M. Lima , A. Lowitz , G. Mahler , A. Mantz , J. L. Marshall , M. McDonald , J. J. McMahon , J. Mena-Fernández , S. S. Meyer , R. Miquel , J. Montgomery , T. Natoli , J. P. Nibarger , G. I. Noble , V. Novosad , R. L. C. Ogando , S. Padin , P. Paschos , S. Patil , A. A. Plazas Malagón , C. Pryke , C. L. Reichardt , J. Roberson , A. K. Romer , C. Romero , J. E. Ruhl , B. R. Saliwanchik , L. Salvati , S. Samuroff , E. Sanchez , B. Santiago , A. Sarkar , A. Saro , K. K. Schaffer , K. Sharon , C. Sievers , G. Smecher , M. Smith , T. Somboonpanyakul , M. Sommer , B. Stalder , A. A. Stark , J. Stephen , V. Strazzullo , E. Suchyta , M. E. C. Swanson , G. Tarle , D. Thomas , C. Tucker , D. L. Tucker , T. Veach , J. D. Vieira , A. von der Linden , G. Wang , N. Whitehorn , W. L. K. Wu , V. Yefremenko , M. Young , J. A. Zebrowski , H. Zohren , DES Collaboration , SPT Collaboration

The ability to precisely quantify similarity between various entities has been a fundamental complication in various problem spaces specifically in the classification of cellular images. Contemporary similarity measures applied in the…

Computer Vision and Pattern Recognition · Computer Science 2018-12-04 D Yoan L. Mekontchou Yomba

Diverse classes of proteins function through large-scale conformational changes; sophisticated enhanced sampling methods have been proposed to generate these macromolecular transition paths. As such paths are curves in a high-dimensional…

Quantitative Methods · Quantitative Biology 2015-10-27 Sean L. Seyler , Avishek Kumar , Michael F. Thorpe , Oliver Beckstein

We use measurements from the South Pole Telescope (SPT) Sunyaev Zel'dovich (SZ) cluster survey in combination with X-ray measurements to constrain cosmological parameters. We present a statistical method that fits for the scaling relations…

Propensity Score Matching (PSM) stands as a widely embraced method in comparative effectiveness research. PSM crafts matched datasets, mimicking some attributes of randomized designs, from observational data. In a valid PSM design where all…

Methodology · Statistics 2024-11-15 Fei Wan
‹ Prev 1 2 3 10 Next ›