Related papers: Efficient fidelity control by stepwise nucleotide …
In this era of big data, feature selection techniques, which have long been proven to simplify the model, makes the model more comprehensible, speed up the process of learning, have become more and more important. Among many developed…
Non-classical two-step nucleation including preordering and crystal nucleation has been widely proposed to challenge the one-step nucleation framework in diverse materials, while what drives preordering has not been explicitly resolved yet.…
The emergence of functional oligonucleotides on early Earth required a molecular selection mechanism to screen for specific sequences with prebiotic functions. Cyclic processes such as daily temperature oscillations were ubiquitous in this…
Generative models have the ability to synthesize data points drawn from the data distribution, however, not all generated samples are high quality. In this paper, we propose using a combination of coresets selection methods and ``entropic…
It is a long-standing question in origin-of-life research whether the information content of replicating molecules can be maintained in the presence of replication errors. Extending standard quasispecies models of non-enzymatic replication,…
The integration time step is a critical determinant of performance in molecular dynamics simulations, governing the trade-off between speed and fidelity. Although 2 fs remains the standard in atomistic biomolecular simulations, the push for…
At nanometer manufacturing technology nodes, process variations significantly affect circuit performance. To combat them, post- silicon clock tuning buffers can be deployed to balance timing bud- gets of critical paths for each individual…
In this paper, we present an efficient way of performing stepwise selection in the class of decomposable models. The main contribution of the paper is a simple characterization of the edges that canbe added to a decomposable model while…
Emergence and maintenance of polymers with complex sequences is a major question in the study of origins of life. To answer this, we studied a model polymerization reaction, where polymers are synthesized by stepwise ligation from two types…
Despite the extensive literature on training loss functions, the evaluation of generalization on the validation set remains underexplored. In this work, we conduct a systematic empirical and statistical study of how the validation criterion…
Protein folding is the intricate process by which a linear sequence of amino acids self-assembles into a unique three-dimensional structure. Protein folding kinetics is the study of pathways and time-dependent mechanisms a protein undergoes…
Feature selection aims to identify the most pattern-discriminative feature subset. In prior literature, filter (e.g., backward elimination) and embedded (e.g., Lasso) methods have hyperparameters (e.g., top-K, score thresholding) and tie to…
The assembly of proteins in membranes plays a key role in many crucial cellular pathways. Despite their importance, characterizing transmembrane assembly remains challenging for experiments and simulations. Equilibrium molecular dynamics…
Early fault detection and fault prognosis are crucial to ensure efficient and safe operations of complex engineering systems such as the Spallation Neutron Source (SNS) and its power electronics (high voltage converter modulators).…
Evolving genomes increase a number of their genes by gene duplications. To escape degradation in a functionless pseudogene, any gene duplicate needs to be guarded by negative (purifying) selection from otherwise inevitable fixation of…
Linear attention has attracted interest as a computationally efficient approximation to softmax attention, especially for long sequences. Recent studies have explored distilling softmax attention in pre-trained Transformers into linear…
The determination of a patient's DNA sequence can, in principle, reveal an increased risk to fall ill with particular diseases [1,2] and help to design "personalized medicine" [3]. Moreover, statistical studies and comparison of genomes [4]…
In recent years, deep learning has been at the center of analytics due to its impressive empirical success in analyzing complex data objects. Despite this success, most of the existing tools behave like black-box machines, thus the…
Genomics biobanks are information treasure troves with thousands of phenotypes (e.g., diseases, traits) and millions of single nucleotide polymorphisms (SNPs). The development of methodologies that provide reproducible discoveries is…
A hydrodynamic model for determining the electrophoretic speed of a polyelectrolyte through a nanopore is presented. It is assumed that the speed is determined by a balance of electrical and viscous forces arising from within the pore and…