English
Related papers

Related papers: Entropic Approach for Reduction of Amino Acid Alph…

200 papers

The primitive data for deducing the Miyazawa-Jernigan contact energy or BLOSUM score matrix consists of pair frequency counts. Each amino acid corresponds to a conditional probability distribution. Based on the deviation of such conditional…

Biological Physics · Physics 2009-11-07 Xin Liu , Di Liu , Ji Qi , Wei-Mou Zheng

Window profiles of amino acids in protein sequences are taken as a description of the amino acid environment. The relative entropy or Kullback-Leibler distance derived from profiles is used as a measure of dissimilarity for comparison of…

Biological Physics · Physics 2009-11-07 Xin Liu , Li-mei Zhang , Shan Guan , Wei-Mou Zheng

We present and implement a distance-based clustering of amino acids within the framework of a statistically derived interaction matrix and show that the resulting groups faithfully reproduce, for well-designed sequences, thermodynamic…

Statistical Mechanics · Physics 2009-10-31 Marek Cieplak , Neal S. Holter , Amos Maritan , Jayanth R. Banavar

We introduce a tensor-based clustering method to extract sparse, low-dimensional structure from high-dimensional, multi-indexed datasets. This framework is designed to enable detection of clusters of data in the presence of structural…

Quantitative Methods · Quantitative Biology 2019-02-11 Anna Seigal , Mariano Beguerisse-Díaz , Birgit Schoeberl , Mario Niepel , Heather A. Harrington

In the present work, we review the fundamental methods which have been developed in the last few years for classifying into families and clans the distribution of amino acids in protein databases. This is done through functions of random…

Biomolecules · Quantitative Biology 2018-06-15 R. P. Mondaini , S. C. de Albuquerque Neto

We present a clustering method and provide a theoretical analysis and an explanation to a phenomenon encountered in the applied statistical literature since the 1990's. This phenomenon is the natural adaptability of the order when using a…

Statistics Theory · Mathematics 2022-03-23 Thierry Dumont

In a statistical approach to protein structure analysis, Miyazawa and Jernigan (MJ) derived a $20\times 20$ matrix of inter-residue contact energies between different types of amino acids. Using the method of eigenvalue decomposition, we…

Statistical Mechanics · Physics 2009-10-28 Hao Li , Chao Tang , Ned Wingreen

In this paper, we address the problem of identifying protein functionality using the information contained in its aminoacid sequence. We propose a method to define sequence similarity relationships that can be used as input for…

Applications · Statistics 2007-11-12 A. G. Flesia , R. Fraiman , F. G. Leonardi

We study clustering methods for binary data, first defining aggregation criteria that measure the compactness of clusters. Five new and original methods are introduced, using neighborhoods and population behavior combinatorial optimization…

A method is described where the aminoacyl-tRNA synthetase system is used to create very small devices for quantitative analysis of the amino acids that occur in proteins. The basis of the method is that each of the 20 synthetases and/or a…

Biological Physics · Physics 2007-05-23 Edward Shipwash

What are proteins made from, as the working parts of the living cells protein machines? To answer this question, we need a technology to disassemble proteins onto elementary func-tional details and to prepare lumped description of such…

Biomolecules · Quantitative Biology 2007-11-05 A. N. Gorban , M. Kudryashev , T. Popova

We attempt to set a mathematical foundation of immunology and amino acid chains. To measure the similarities of these chains, a kernel on strings is defined using only the sequence of the chains and a good amino acid substitution matrix…

Machine Learning · Statistics 2012-06-26 Wen-Jun Shen , Hau-San Wong , Quan-Wu Xiao , Xin Guo , Stephen Smale

A novel methodology is proposed for clustering multivariate time series data using energy distance defined in Sz\'ekely and Rizzo (2013). Specifically, a dissimilarity matrix is formed using the energy distance statistic to measure…

Methodology · Statistics 2024-03-13 Richard A. Davis , Leon Fernandes , Konstantinos Fokianos

We define a general notion of entropy in elementary, algebraic terms. Based on that, weak forms of a scalar product and a distance measure are derived. We give basic properties of these quantities, generalize the Cauchy-Schwarz inequality,…

Spectral Theory · Mathematics 2024-04-10 Martin Schlather

An important issue in clustering concerns the avoidance of false positives while searching for clusters. This work addressed this problem considering agglomerative methods, namely single, average, median, complete, centroid and Ward's…

Machine Learning · Computer Science 2020-06-30 Eric K. Tokuda , Cesar H. Comin , Luciano da F. Costa

The possibility of deriving the contact potentials between amino acids from their frequencies of occurence in proteins is discussed in evolutionary terms. This approach allows the use of traditional thermodynamics to describe such…

Biomolecules · Quantitative Biology 2009-11-10 G. Tiana , M. Colombo , D. Provasi , R. A. Broglia

This paper explores the problem of clustering ensemble, which aims to combine multiple base clusterings to produce better performance than that of the individual one. The existing clustering ensemble methods generally construct a…

Machine Learning · Computer Science 2020-12-17 Yuheng Jia , Hui Liu , Junhui Hou , Qingfu Zhang

We perform an exhaustive analysis of genome statistics for organisms, particularly extremophiles, growing in a wide range of physicochemical conditions. Specifically, we demonstrate how the correlation between the frequency of amino acids…

Genomics · Quantitative Biology 2013-09-19 Benjamin Greenbaum , Pradeep Kumar , Albert Libchaber

Clustering methods with dimension reduction have been receiving considerable wide interest in statistics lately and a lot of methods to simultaneously perform clustering and dimension reduction have been proposed. This work presents a novel…

Methodology · Statistics 2014-06-17 Michio Yamamoto , Kenichi Hayashi

Clustering of mixed-type datasets can be a particularly challenging task as it requires taking into account the associations between variables with different level of measurement, i.e., nominal, ordinal and/or interval. In some cases,…

Methodology · Statistics 2022-04-22 Odysseas Moschidis , Angelos Markos , Theodore Chadjipadelis
‹ Prev 1 2 3 10 Next ›