English
Related papers

Related papers: A statistical physics perspective on alignment-ind…

200 papers

Thermostability is an important prerequisite for enzymes employed for industrial applications. Several machine learning based models have thus been formulated for protein classification based on this particular trait. These models have…

Quantitative Methods · Quantitative Biology 2021-03-08 Jithin S. Sunny , Lilly M. Saleena

Proteins created by combinatorial methods in vitro are an important source of information for understanding sequence-structure-function relationships. Alignments of folded proteins from combinatorial libraries can be analyzed using methods…

Biomolecules · Quantitative Biology 2007-05-23 Jeffrey B. Endelman , Jesse D. Bloom , Christopher R. Otey , Marco Landwehr , Frances H. Arnold

Bayesian likelihood-free methods implement Bayesian inference using simulation of data from the model to substitute for intractable likelihood evaluations. Most likelihood-free inference methods replace the full data set with a summary…

Methodology · Statistics 2020-10-16 Yinan Mao , Xueou Wang , David J. Nott , Michael Evans

Alignment-free sequence analysis approaches provide important alternatives over multiple sequence alignment (MSA) in biological sequence analysis because alignment-free approaches have low computation complexity and are not dependent on…

Computational Engineering, Finance, and Science · Computer Science 2016-08-02 Changchuan Yin , Xuemeng E. Yin , Jiasong Wang

Sequence comparison is a widely used computational technique in modern molecular biology. In spite of the frequent use of sequence comparisons the important problem of assigning statistical significance to a given degree of similarity is…

Quantitative Methods · Quantitative Biology 2007-05-23 Ralf Bundschuh , Nicholas Chia

Physical theories that depend on many parameters or are tested against data from many different experiments pose unique challenges to statistical inference. Many models in particle physics, astrophysics and cosmology fall into one or both…

We present a statistical mechanics approach to the protein folding problem. We first review some of the basic properties of proteins, and introduce some physical models to describe their thermodynamics. These models rely on a random…

Disordered Systems and Neural Networks · Physics 2008-02-03 T. Garel , H. Orland , E. Pitard

The frequencies of A, C, G and T in mitochondrial DNA vary among species due to unequal rates of mutation between the bases. The frequencies of bases at four-fold degenerate sites respond directly to mutation pressure. At 1st and 2nd…

Populations and Evolution · Quantitative Biology 2016-09-08 Daniel Urbina , Bin Tang , Paul G. Higgs

Motivation: Alignment-free distance and similarity functions (AF functions, for short) are a well established alternative to two and multiple sequence alignments for many genomic, metagenomic and epigenomic tasks. Due to data-intensive…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-10-26 Umberto Ferraro Petrillo , Francesco Palini , Giuseppe Cattaneo , Raffaele Giancarlo

The sequence of a protein is not only constrained by its physical and biochemical properties under current selection, but also by features of its past evolutionary history. Understanding the extent and the form that these evolutionary…

Populations and Evolution · Quantitative Biology 2015-06-22 Mathieu Hemery , Olivier Rivoire

Eliciting preferences from human judgements is inherently imprecise, yet most decision analysis methods force a single priority vector from pairwise comparisons, discarding the information embedded in inconsistencies. We instead leverage…

General Economics · Economics 2026-02-27 Salvatore Greco , Sajid Siraj , Michele Lundy

Efficient automatic protein classification is of central importance in genomic annotation. As an independent way to check the reliability of the classification, we propose a statistical approach to test if two sets of protein domain…

Atomic-level simulations are widely used to study biomolecules and their dynamics. A common goal in such studies is to compare simulations of a molecular system under several conditions -- for example, with various mutations or bound…

Biomolecules · Quantitative Biology 2025-01-07 Martin Vögele , Neil J. Thomson , Sang T. Truong , Jasper McAvity , Ulrich Zachariae , Ron O. Dror

A deep neural network based architecture was constructed to predict amino acid side chain conformation with unprecedented accuracy. Amino acid side chain conformation prediction is essential for protein homology modeling and protein design.…

Biomolecules · Quantitative Biology 2017-07-27 Ke Liu , Xiangyan Sun , Jun Ma , Zhenyu Zhou , Qilin Dong , Shengwen Peng , Junqiu Wu , Suocheng Tan , Günter Blobel , Jie Fan

Simple hidden Markov models are proposed for predicting secondary structure of a protein from its amino acid sequence. Since the length of protein conformation segments varies in a narrow range, we ignore the duration effect of length…

Biological Physics · Physics 2007-05-23 Wei-Mou Zheng

Circular permutation connects the N and C termini of a protein and concurrently cleaves elsewhere in the chain, providing an important mechanism for generating novel protein fold and functions. However, their in genomes is unknown because…

Biomolecules · Quantitative Biology 2016-11-17 T. Andrew Binkowski , Bhaskar DasGupta , Jie Liang

We present a new distribution-free conformal prediction algorithm for sequential data (e.g., time series), called the \textit{sequential predictive conformal inference} (\texttt{SPCI}). We specifically account for the nature that time…

Machine Learning · Statistics 2023-05-31 Chen Xu , Yao Xie

To infer the parameters of mechanistic models with intractable likelihoods, techniques such as approximate Bayesian computation (ABC) are increasingly being adopted. One of the main disadvantages of ABC in practical situations, however, is…

Computation · Statistics 2018-08-03 Jonathan U Harrison , Ruth E Baker

In this paper, we address the problem of identifying protein functionality using the information contained in its aminoacid sequence. We propose a method to define sequence similarity relationships that can be used as input for…

Applications · Statistics 2007-11-12 A. G. Flesia , R. Fraiman , F. G. Leonardi

Distance-based regression model, as a nonparametric multivariate method, has been widely used to detect the association between variations in a distance or dissimilarity matrix for outcomes and predictor variables of interest in genetic…

Statistics Theory · Mathematics 2022-03-14 Yuke Shi , Wei Zhang , Aiyi Liu , Qizhai Li