English
Related papers

Related papers: Seven clusters in genomic triplet distributions

200 papers

Coding information is the main source of heterogeneity (non-randomness) in the sequences of bacterial genomes. This information can be naturally modeled by analysing cluster structures in the "in-phase" triplet distributions of relatively…

Genomics · Quantitative Biology 2011-11-09 A. N. Gorban , T. G. Popova , A. Yu. Zinovyev

Previously, a seven-cluster pattern claiming to be a universal one in bacterial genomes has been reported. Keeping in mind the most popular theory of chloroplast origin, we checked whether a similar pattern is observed in chloroplast…

Genomics · Quantitative Biology 2018-02-09 Michael Sadovsky , Maria Senashova , Andrew Malyshev

An approach based on using the idea of distinguished coding phase in explicit form for identification of protein-coding regions (exons) in whole genome has been proposed. For several genomes an optimal window length for averaging GC-content…

Biological Physics · Physics 2007-05-23 Alexander N. Gorban , Andrey Yu. Zinovyev , Tatyana G. Popova

We studied the structuredness in a chloroplast genome of Siberian larch. The clusters in 63-dimensional space were identified with elastic map technique, where the objects to be clusterized are the different fragments of the genome. A…

Genomics · Quantitative Biology 2016-04-18 Michael G. Sadovsky , Eugenia I. Bondar , Yuliya A. Putintseva , Konstantin V. Krutovsky

We studied a relations between the triplet frequency composition of mitochondria genomes, and the phylogeny of their bearers. First, the clusters in 63dimensional space were developed due to $K$-means. Second, the clade composition of those…

Genomics · Quantitative Biology 2014-05-21 Michael G. Sadovsky

We propose an algorithm for clustering high dimensional data. If $P$ features for $N$ objects are represented in an $N\times P$ matrix ${\bf X}$, where $N\ll P$, the method is based on exploiting the cluster-dependent structure of the…

Machine Learning · Statistics 2018-11-05 Shahina Rahman , Valen E. Johnson

A general method to obtain a representation of the structural landscape of nanoparticles in terms of a limited number of variables is proposed. The method is applied to a large dataset of parallel tempering molecular dynamics simulations of…

Clustering is typically measured by the ratio of triangles to all triples, open or closed. Generating clustered networks, and how clustering affects dynamics on networks, is reasonably well understood for certain classes of networks…

Physics and Society · Physics 2014-10-22 Martin Ritchie , Luc Berthouze , Thomas House , Istvan Z. Kiss

We applied the newly developed rose diagram overlay method to detect the layered structure of 88 nearby open clusters ($\leq$500~pc) on the three projections after the distance correction of their member stars, based on the catalog in…

Astrophysics of Galaxies · Physics 2023-05-18 Qingshun Hu , Yu Zhang , Ali Esamdin , Hong Wang , Mingfeng Qin

Patchwork learning arises as a new and challenging data collection paradigm where both samples and features are observed in fragmented subsets. Due to technological limits, measurement expense, or multimodal data integration, such patchwork…

Methodology · Statistics 2024-06-21 Lili Zheng , Andersen Chang , Genevera I. Allen

In this paper we propose a unified framework to simultaneously discover the number of clusters and group the data points into them using subspace clustering. Real data distributed in a high-dimensional space can be disentangled into a union…

Computer Vision and Pattern Recognition · Computer Science 2019-07-24 Jie Liang , Jufeng Yang , Ming-Ming Cheng , Paul L. Rosin , Liang Wang

In this work we seek clusters of genomic words in human DNA by studying their inter-word lag distributions. Due to the particularly spiked nature of these histograms, a clustering procedure is proposed that first decomposes each…

Applications · Statistics 2021-01-13 Ana Helena Tavares , Jakob Raymaekers , Peter J. Rousseeuw , Paula Brito , Vera Afreixo

Protein representation learning is a challenging task that aims to capture the structure and function of proteins from their amino acid sequences. Previous methods largely ignored the fact that not all amino acids are equally important for…

Machine Learning · Computer Science 2024-04-02 Ruijie Quan , Wenguan Wang , Fan Ma , Hehe Fan , Yi Yang

The expression levels of many thousands of genes can be measured simultaneously by DNA microarrays (chips). This novel experimental tool has revolutionized research in molecular biology and generated considerable excitement. A typical…

Biological Physics · Physics 2007-05-23 Eytan Domany

Analyses of targeted genomic sequencing data from next-generation-sequencing (NGS) technologies typically involves mapping reads to a reference sequence or clustering reads. For a number of species a reference genome is not available so the…

Genomics · Quantitative Biology 2016-02-16 Raunaq Malhotra , Daniel Elleder , Le Bao , David R Hunter , Raj Acharya , Mary Poss

Several modern genomic technologies, such as DNA-Methylation arrays, measure spatially registered probes that number in the hundreds of thousands across multiplechromosomes. The measured probes are by themselves less interesting…

Applications · Statistics 2016-11-16 John Nagorski , Genevera I. Allen

We propose a novel methodology for feature screening in clustering massive datasets, in which both the number of features and the number of observations can potentially be very large. Taking advantage of a fusion penalization based convex…

Methodology · Statistics 2017-10-05 Trambak Banerjee , Gourab Mukherjee , Peter Radchenko

Discrete mixture models provide a well-known basis for effective clustering algorithms, although technical challenges have limited their scope. In the context of gene-expression data analysis, a model is presented that mixes over a finite…

Methodology · Statistics 2012-11-12 Michael A. Newton , Lisa M. Chung

We propose a novel method to cluster gene networks. Based on a dissimilarity built using correlation structures, we consider networks that connect all the genes based on the strength of their dissimilarity. The large number of genes require…

Statistics Theory · Mathematics 2016-07-07 A-C Brunet , J-M Azais , J-M Loubes , J Amar , R Burcelin

In cluster tomography, we propose measuring the number of clusters $N$ intersected by a line segment of length $\ell$ across a finite sample. As expected, the leading order of $N(\ell)$ scales as $a\ell$, where $a$ depends on microscopic…

Disordered Systems and Neural Networks · Physics 2024-02-13 Helen S. Ansell , Samuel J. Frank , István A. Kovács
‹ Prev 1 2 3 10 Next ›