中文
相关论文

相关论文: Compression ratios based on the Universal Similari…

200 篇论文

The Universal Similarity Metric (USM) has been demonstrated to give practically useful measures of "similarity" between sequence data. Here we have used the USM as an alternative distance metric in a K-Nearest Neighbours (K-NN) learner to…

机器学习 · 计算机科学 2024-05-13 David Lindsay , Sian Lindsay

We present a new method for clustering based on compression. The method doesn't use subject-specific features or background knowledge, and works as follows: First, we determine a universal similarity distance, the normalized compression…

计算机视觉与模式识别 · 计算机科学 2007-05-23 Rudi Cilibrasi , Paul Vitanyi

We present a new similarity measure based on information theoretic measures which is superior than Normalized Compression Distance for clustering problems and inherits the useful properties of conditional Kolmogorov complexity. We show that…

机器学习 · 统计学 2014-10-22 Andrey Bogomolov , Bruno Lepri , Fabio Pianesi

A new class of distances appropriate for measuring similarity relations between sequences, say one type of similarity per distance, is studied. We propose a new ``normalized information distance'', based on the noncomputable notion of…

计算复杂性 · 计算机科学 2011-11-09 Ming Li , Xin Chen , Xin Li , Bin Ma , Paul Vitanyi

Enormous volumes of short reads data from next-generation sequencing (NGS) technologies have posed new challenges to the area of genomic sequence comparison. The multiple sequence alignment approach is hardly applicable to NGS data due to…

基因组学 · 定量生物学 2020-03-25 Ngoc Hieu Tran , Xin Chen

Distance-based clustering and classification are widely used in various fields to group mixed numeric and categorical data. In many algorithms, a predefined distance measurement is used to cluster data points based on their dissimilarity.…

机器学习 · 计算机科学 2024-10-14 Jesse S. Ghashti , John R. J. Thompson

The advent of high dimensional single cell data in the biomedical sciences has necessitated the development of dimensionality-reduction tools. t-SNE and UMAP are the two most frequently used approaches, allowing clear visualisation of…

We study the problem of efficiently clustering protein sequences in a limited information setting. We assume that we do not know the distances between the sequences in advance, and must query them during the execution of the algorithm. Our…

数据结构与算法 · 计算机科学 2015-03-17 Konstantin Voevodski , Maria-Florina Balcan , Heiko Roglin , Shang-Hua Teng , Yu Xia

A multitude of measures have been proposed to quantify the similarity between protein 3-D structure. Among these measures, contact map overlap (CMO) maximization deserved sustained attention during past decade because it offers a fine…

定量方法 · 定量生物学 2009-04-20 Rumen Andonov , Nicola Yanev , Noël Malod-Dognin

Single-cell omics enable the profiles of cells, which contain large numbers of biological features, to be quantified. Cluster analysis, a dimensionality reduction process, is used to reduce the dimensions of the data to make it…

基因组学 · 定量生物学 2024-06-06 Okezue Bell , Arthur Lee , Elizabeth Engle

We study the problem of how well a tree metric is able to preserve the sum of pairwise distances of an arbitrary metric. This problem is closely related to low-stretch metric embeddings and is interesting by its own flavor from the line of…

数据结构与算法 · 计算机科学 2013-01-16 Mong-Jen Kao , Der-Tsai Lee , Dorothea Wagner

Normalized compression distance (NCD) is a parameter-free, feature-free, alignment-free, similarity measure between a pair of finite objects based on compression. However, it is not sufficient for all applications. We propose an NCD of…

计算机视觉与模式识别 · 计算机科学 2016-01-28 Andrew R. Cohen , Paul M. B. Vitanyi

Research into the classification of time series has made enormous progress in the last decade. The UCR time series archive has played a significant role in challenging and guiding the development of new learners for time series…

We consider the problem of metric learning subject to a set of constraints on relative-distance comparisons between the data items. Such constraints are meant to reflect side-information that is not expressed directly in the feature vectors…

机器学习 · 计算机科学 2016-12-06 Ehsan Amid , Aristides Gionis , Antti Ukkonen

In another related work, U-statistics were used for non-asymptotic "average-case" analysis of random compressed sensing matrices. In this companion paper the same analytical tool is adopted differently - here we perform non-asymptotic…

信息论 · 计算机科学 2015-06-11 Fabian Lim , Vladimir Stojanovic

Cosmic shear, galaxy clustering, and the abundance of massive halos each probe the large-scale structure of the Universe in complementary ways. We present cosmological constraints from the joint analysis of the three probes, building on the…

宇宙学与河外天体物理 · 物理学 2025-03-14 S. Bocquet , S. Grandis , E. Krause , C. To , L. E. Bleem , M. Klein , J. J. Mohr , T. Schrabback , A. Alarcon , O. Alves , A. Amon , F. Andrade-Oliveira , E. J. Baxter , K. Bechtol , M. R. Becker , G. M. Bernstein , J. Blazek , H. Camacho , A. Campos , A. Carnero Rosell , M. Carrasco Kind , R. Cawthon , C. Chang , R. Chen , A. Choi , J. Cordero , M. Crocce , C. Davis , J. DeRose , H. T. Diehl , S. Dodelson , C. Doux , A. Drlica-Wagner , K. Eckert , T. F. Eifler , F. Elsner , J. Elvin-Poole , S. Everett , X. Fang , A. Ferté , P. Fosalba , O. Friedrich , J. Frieman , M. Gatti , G. Giannini , D. Gruen , R. A. Gruendl , I. Harrison , W. G. Hartley , K. Herner , H. Huang , E. M. Huff , D. Huterer , M. Jarvis , N. Kuropatkin , P. -F. Leget , P. Lemos , A. R. Liddle , N. MacCrann , J. McCullough , J. Muir , J. Myles , A. Navarro-Alsina , S. Pandey , Y. Park , A. Porredon , J. Prat , M. Raveri , R. P. Rollins , A. Roodman , R. Rosenfeld , E. S. Rykoff , C. Sánchez , J. Sanchez , L. F. Secco , I. Sevilla-Noarbe , E. Sheldon , T. Shin , M. A. Troxel , I. Tutusaus , T. N. Varga , N. Weaverdyck , R. H. Wechsler , H. -Y. Wu , B. Yanny , B. Yin , Y. Zhang , J. Zuntz , T. M. C. Abbott , P. A. R. Ade , M. Aguena , S. Allam , S. W. Allen , A. J. Anderson , B. Ansarinejad , J. E. Austermann , M. Bayliss , J. A. Beall , A. N. Bender , B. A. Benson , F. Bianchini , M. Brodwin , D. Brooks , L. Bryant , D. L. Burke , R. E. A. Canning , J. E. Carlstrom , J. Carretero , F. J. Castander , C. L. Chang , P. Chaubal , H. C. Chiang , T-L. Chou , R. Citron , C. Corbett Moran , M. Costanzi , T. M. Crawford , A. T. Crites , L. N. da Costa , M. E. S. Pereira , T. M. Davis , T. de Haan , M. A. Dobbs , P. Doel , W. Everett , A. Farahi , B. Flaugher , A. M. Flores , B. Floyd , J. Gallicchio , E. Gaztanaga , E. M. George , M. D. Gladders , N. Gupta , G. Gutierrez , N. W. Halverson , S. R. Hinton , J. Hlavacek-Larrondo , G. P. Holder , D. L. Hollowood , W. L. Holzapfel , J. D. Hrubes , N. Huang , J. Hubmayr , K. D. Irwin , D. J. James , F. Kéruzoré , G. Khullar , K. Kim , L. Knox , R. Kraft , K. Kuehn , O. Lahav , A. T. Lee , S. Lee , D. Li , C. Lidman , M. Lima , A. Lowitz , G. Mahler , A. Mantz , J. L. Marshall , M. McDonald , J. J. McMahon , J. Mena-Fernández , S. S. Meyer , R. Miquel , J. Montgomery , T. Natoli , J. P. Nibarger , G. I. Noble , V. Novosad , R. L. C. Ogando , S. Padin , P. Paschos , S. Patil , A. A. Plazas Malagón , C. Pryke , C. L. Reichardt , J. Roberson , A. K. Romer , C. Romero , J. E. Ruhl , B. R. Saliwanchik , L. Salvati , S. Samuroff , E. Sanchez , B. Santiago , A. Sarkar , A. Saro , K. K. Schaffer , K. Sharon , C. Sievers , G. Smecher , M. Smith , T. Somboonpanyakul , M. Sommer , B. Stalder , A. A. Stark , J. Stephen , V. Strazzullo , E. Suchyta , M. E. C. Swanson , G. Tarle , D. Thomas , C. Tucker , D. L. Tucker , T. Veach , J. D. Vieira , A. von der Linden , G. Wang , N. Whitehorn , W. L. K. Wu , V. Yefremenko , M. Young , J. A. Zebrowski , H. Zohren , DES Collaboration , SPT Collaboration

The ability to precisely quantify similarity between various entities has been a fundamental complication in various problem spaces specifically in the classification of cellular images. Contemporary similarity measures applied in the…

计算机视觉与模式识别 · 计算机科学 2018-12-04 D Yoan L. Mekontchou Yomba

Diverse classes of proteins function through large-scale conformational changes; sophisticated enhanced sampling methods have been proposed to generate these macromolecular transition paths. As such paths are curves in a high-dimensional…

定量方法 · 定量生物学 2015-10-27 Sean L. Seyler , Avishek Kumar , Michael F. Thorpe , Oliver Beckstein

Propensity Score Matching (PSM) stands as a widely embraced method in comparative effectiveness research. PSM crafts matched datasets, mimicking some attributes of randomized designs, from observational data. In a valid PSM design where all…

统计方法学 · 统计学 2024-11-15 Fei Wan
‹ 上一页 1 2 3 10 下一页 ›