Related papers: Predicting evolutionary site variability from stru…
This piece serves two purposes. Firstly, it aims at elucidating the role of epistasis in shaping, at a molecular level, the evolutionary paths of proteins, as well as the extent to which these epistatic effects are the outcome of an…
Protein aggregation occurs when misfolded or unfolded proteins physically bind together, and can promote the development of various amyloid diseases. This study aimed to construct surrogate models for predicting protein aggregation via…
Metrics for indirectly predicting the folding rates of RNA sequences are of interest. In this letter, we introduce a simple metric of RNA structural complexity, which accounts for differences in the energetic contributions of RNA base…
Survival analysis concerns the task of predicting the time until an event occurs. Often used in the medical field, survival analysis deals with incomplete (i.e., censored) data, for instance, from patients who did not experience the event…
Recommender Systems (RS) shape the filtering and curation of online content, yet we have limited understanding of how predictable their recommendation outputs are. We propose data-driven metrics that quantify the predictability of…
A complete time-parameterized statistical model quantifying the divergent evolution of protein structures in terms of the patterns of conservation of their secondary structures is inferred from a large collection of protein 3D structure…
Feature models are widely used to capture the configuration space of software systems. Although automated reasoning has been studied for detecting problematic features and supporting configuration tasks, significantly less attention has…
Protein complex formation is a central problem in biology, being involved in most of the cell's processes, and essential for applications, e.g. drug design or protein engineering. We tackle rigid body protein-protein docking, i.e.,…
In the framework of a lattice-model study of protein folding, we investigate the interplay between designability, thermodynamic stability, and kinetics. To be ``protein-like'', heteropolymers must be thermodynamically stable, stable against…
An effective potential function is critical for protein structure prediction and folding simulation. For simplified models of proteins where coordinates of only $C_\alpha$ atoms need to be specified, an accurate potential function is…
In this work we employ various methods of analysis (unfolding simulations and comparative analysis of structures and sequences of proteomes of thermophilic organisms) to show that organisms can follow two major strategies of thermophilic…
Comprehensive discovery of structural variation (SV) in human genomes from DNA sequencing requires the integration of multiple alignment signals including read-pair, split-read and read-depth. However, owing to inherent technical…
The packaging of genetic material within a protein shell, called the capsid, marks a pivotal step in the life cycle of numerous single-stranded RNA viruses. Understanding how hundreds, or even thousands, of proteins assemble around the…
The coding space of protein sequences is shaped by evolutionary constraints set by requirements of function and stability. We show that the coding space of a given protein family--the total number of sequences in that family--can be…
Although density functional theory provides reliable predictions for the static properties of simple fluids under confinement, a theory of comparative accuracy for the transport coefficients has yet to emerge. Nonetheless, there is evidence…
Single domain proteins are thought to be tightly packed. The introduction of voids by mutations is often regarded as destabilizing. In this study we show that packing density for single domain proteins decreases with chain length. We find…
Forecasting the change in the distribution of viral variants is crucial for therapeutic design and disease surveillance. This task poses significant modeling challenges due to the sharp differences in virus distributions across…
High-quality training datasets are crucial for the development of effective protein design models, but existing synthetic datasets often include unfavorable sequence-structure pairs, impairing generative model performance. We leverage…
We study statistical properties of interacting protein-like surfaces and predict two strong, related effects: (i) statistically enhanced self-attraction of proteins; (ii) statistically enhanced attraction of proteins with similar…
Prediction of ligand binding sites of proteins is a fundamental and important task for understanding the function of proteins and screening potential drugs. Most existing methods require experimentally determined protein holo-structures as…