English
Related papers

Related papers: A Dirichlet process mixture of hidden Markov model…

200 papers

Protein inverse folding is a fundamental problem in bioinformatics, aiming to recover the amino acid sequences from a given protein backbone structure. Despite the success of existing methods, they struggle to fully capture the intricate…

Machine Learning · Computer Science 2024-12-13 Chenglin Wang , Yucheng Zhou , Zijie Zhai , Jianbing Shen , Kai Zhang

Classification of proteins based on their structure provides a valuable resource for studying protein structure, function and evolutionary relationships. With the rapidly increasing number of known protein structures, manual and…

Computational Engineering, Finance, and Science · Computer Science 2009-07-14 Oktie Hassanzadeh

We propose that protein loops can be interpreted as topological domain-wall solitons. They interpolate between ground states that are the secondary structures like alpha-helices and beta-strands. Entire proteins can then be folded simply by…

Biological Physics · Physics 2014-11-20 M. N. Chernodub , Shuangwei Hu , Antti J. Niemi

Variation in the evolutionary process across the sites of nucleotide sequence alignments is well established, and is an increasingly pervasive feature of datasets composed of gene regions sampled from multiple loci and/or different genomes.…

Populations and Evolution · Quantitative Biology 2014-09-04 Brian R. Moore , Jim McGuire , Fredrik Ronquist , John P. Huelsenbeck

We introduce a formulation for normal mode analyses of globular proteins that significantly improves on an earlier, 1-parameter formulation (M. Tirion, PRL 77, 1905 (1996)) that characterized the slow modes associated with protein data bank…

Biological Physics · Physics 2015-04-01 Monique M. Tirion , Daniel ben-Avraham

An important problem in shape analysis is to match configurations of points in space filtering out some geometrical transformation. In this paper we introduce hierarchical models for such tasks, in which the points in the configurations are…

Statistics Theory · Mathematics 2010-03-23 Peter J. Green , Kanti Mardia

This paper gives a method for computing distributions associated with patterns in the state sequence of a hidden Markov model, conditional on observing all or part of the observation sequence. Probabilities are computed for very general…

Methodology · Statistics 2007-12-18 John A. D. Aston , Donald E. K. Martin

Inverse protein folding is a fundamental task in computational protein design, which aims to design protein sequences that fold into the desired backbone structures. While the development of machine learning algorithms for this task has…

Machine Learning · Computer Science 2024-11-05 Yiheng Zhu , Jialu Wu , Qiuyi Li , Jiahuan Yan , Mingze Yin , Wei Wu , Mingyang Li , Jieping Ye , Zheng Wang , Jian Wu

We consider a discrete latent variable model for two-way data arrays, which allows one to simultaneously produce clusters along one of the data dimensions (e.g. exchangeable observational units or features) and contiguous groups, or…

This paper is concerned with statistical methods for the segmental classification of linear sequence data where the task is to segment and classify the data according to an underlying hidden discrete state sequence. Such analysis is…

Methodology · Statistics 2015-05-05 Christopher Yau , Christopher C. Holmes

Learning from 3D protein structures has gained wide interest in protein modeling and structural bioinformatics. Unfortunately, the number of available structures is orders of magnitude lower than the training data sizes commonly used in…

Biomolecules · Quantitative Biology 2022-06-01 Pedro Hermosilla , Timo Ropinski

This paper describes the adaptation of a well-scaling parallel algorithm for computing Morse-Smale segmentations based on path compression to a distributed computational setting. Additionally, we extend the algorithm to efficiently compute…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-09-09 Michael Will , Jonas Lukasczyk , Julien Tierny , Christoph Garth

When modeling the distribution of a set of data by a mixture of Gaussians, there are two possibilities: i) the classical one is using a set of parameters which are the proportions, the means and the variances; ii) the second is to consider…

Data Analysis, Statistics and Probability · Physics 2009-11-13 Ali Mohammad-Djafari

Understanding the protein folding process is an outstanding issue in biophysics; recent developments in molecular dynamics simulation have provided insights into this phenomenon. However, the large freedom of atomic motion hinders the…

Computational Physics · Physics 2020-06-18 Takashi Ichinomiya , Ippei Obayashi , Yasuaki Hiraoka

We propose a network structure discovery model for continuous observations that generalizes linear causal models by incorporating a Gaussian process (GP) prior on a network-independent component, and random sparsity and weight matrices as…

Machine Learning · Computer Science 2017-03-01 Amir Dezfouli , Edwin V. Bonilla , Richard Nock

In recent years, there has been a surge in the development of 3D structure-based pre-trained protein models, representing a significant advancement over pre-trained protein language models in various downstream tasks. However, most existing…

Machine Learning · Computer Science 2024-06-04 Jiale Zhao , Wanru Zhuang , Jia Song , Yaqi Li , Shuqi Lu

Predicting the binding structure of a small molecule ligand to a protein -- a task known as molecular docking -- is critical to drug design. Recent deep learning methods that treat docking as a regression problem have decreased runtime…

Biomolecules · Quantitative Biology 2023-02-14 Gabriele Corso , Hannes Stärk , Bowen Jing , Regina Barzilay , Tommi Jaakkola

Numerous cellular functions rely on protein$\unicode{x2013}$protein interactions. Efforts to comprehensively characterize them remain challenged however by the diversity of molecular recognition mechanisms employed within the proteome. Deep…

Biomolecules · Quantitative Biology 2023-12-08 Julia R. Rogers , Gergő Nikolényi , Mohammed AlQuraishi

Sequence set is a widely-used type of data source in a large variety of fields. A typical example is protein structure prediction, which takes an multiple sequence alignment (MSA) as input and aims to infer structural information from it.…

Biomolecules · Quantitative Biology 2019-06-27 Fusong Ju , Jianwei Zhu , Guozheng Wei , Qi Zhang , Shiwei Sun , Dongbo Bu

This paper proposes a new mathematical approach to characterize native protein structures based on the discrete differential geometry of tetrahedron tiles. In the approach, local structure of proteins is classified into finite types…

Biomolecules · Quantitative Biology 2007-05-23 Naoto Morikawa