Related papers: Mutation model for nucleotide sequences based on c…
This paper presents the Ensemble Nucleotide Byte-level Encoder-Decoder (ENBED) foundation model, analyzing DNA sequences at byte-level precision with an encoder-decoder Transformer architecture. ENBED uses a sub-quadratic implementation of…
Let $q$ be a positive integer and $\mathcal{S}=\left\{x_0,x_1,\ldots,x_{T-1}\right\}\subseteq\mathbb{Z}_q=\{0,1,\ldots,q-1\}$ with $$0\leq x_0<x_1<\ldots< x_{T-1}\leq q-1.$$ We derive from $\mathcal{S}$ three (finite) sequences. 1. For an…
Power spectra of human DNA base C+G frequency distribution in all available contiguous sections exhibit the universal inverse power law form of the statistical normal distribution for the 24 chromosomes. Inverse power law form for power…
This work proposes a markovian memoryless model for the DNA that simplifies enormously the complexity of it. We encode nucleotide sequences into symbolic sequences, called words, from which we establish meaningful length of words and group…
DNA sequencing is the process of determining the exact order of the nucleotide bases of an individual's genome in order to catalogue sequence variation and understand its biological implications. Whole-genome sequencing techniques produce…
Experiments at Jefferson Laboratory, MIT-Bates, LEGS, Mainz, Bonn, GRAAL, and Spring-8 offer new opportunities to understand in detail how nucleon resonance ($N^*$) properties emerge from the nonperturbative aspects of QCD. Preliminary data…
Recent studies of DNA sequence of letters A, C, G and T exhibit the inverse power law form. Inverse power-law form of the power spectra of fractal space-time fluctuations is generic to the dynamical systems in nature and is identified as…
We show how single-molecule unzipping experiments can provide strong evidence that the zero-force melting transition of long molecules of natural dsDNA should be classified as a phase transition of the higher-order type (continuous). We…
An RNA sequence is a word over an alphabet on four elements $\{A,C,G,U\}$ called bases. RNA sequences fold into secondary structures where some bases match one another while others remain unpaired. Pseudoknot-free secondary structures can…
The flexibility of short DNA chains is investigated via computation of the average correlation function between dimers which defines the persistence length. Path integration techniques have been applied to confine the phase space available…
In shotgun sequencing, the input string (typically, a long DNA sequence composed of nucleotide bases) is sequenced as multiple overlapping fragments of much shorter lengths (called \textit{reads}). Modelling the shotgun sequencing pipeline…
The dependence of the frequency distributions due to multiple meanings of words in a text is investigated by deleting letters. By coding the words with fewer letters the number of meanings per coded word increases. This increase is measured…
In many situations, the gene expression signature is a unique marker of the biological state. We study the modification of the gene expression distribution function when the biological state of a system experiences a change. This change may…
Valence double parton distribution functions of the nucleon are evaluated in the framework of a simple model, where the conservation of the longitudinal momentum is taken into account. The leading-order DGLAP QCD evolution from the low…
We consider a novel approach of measuring the homology of DNA sequences based of the variety of optimal alignments in the longest common subsequence sense. The proposed approach is compared with BLAST in measuring the homology of four…
In this paper, we review the literature on statistical long-range correlation in DNA sequences. We examine the current evidence for these correlations, and conclude that a mixture of many length scales (including some relatively long ones)…
The study of correlation structure in the primary sequences of DNA is reviewed. The issues reviewed include: symmetries among 16 base-base correlation functions, accurate estimation of correlation measures, the relationship between $1/f$…
Label distribution learning (LDL) provides a framework wherein a distribution over categories rather than a single category is predicted, with the aim of addressing ambiguity in labeled data. Existing research on LDL mainly focuses on the…
In this short communication, we shall explore a nonlinear discrete dynamical system that naturally occurs in population systems to describe a transmission of a trait from parents to their offspring. We consider a Mendelian inheritance for a…
Convolutional codes are error-correcting linear codes that utilize shift registers to encode. These codes have an arbitrary block size and they can incorporate both past and current information bits. DNA codes represent DNA sequences and…