Related papers: Two-Parameter Characterization of Chromosome-Scale…
This paper presents an approach to reducing the number of fundamental parameters in the Standard Model (SM) using genetic programming, a machine learning technique based on evolutionary algorithms. We outline the core principles of our…
Current popular methods in literature of RNA sequencing normalisation do not account for gene length when compared across samples, whilst adjusting for count biases in the data. This creates a gap in the normalisation as bigger genes in RNA…
Parameter Recombination (PR) methods aim to efficiently compose the weights of a neural network for applications like Parameter-Efficient FineTuning (PEFT) and Model Compression (MC), among others. Most methods typically focus on one…
Genetic algorithms are a well-known method for tackling the problem of variable selection. As they are non-parametric and can use a large variety of fitness functions, they are well-suited as a variable selection wrapper that can be applied…
A recolouring sequence, between $k$-colourings $\alpha$ and $\beta$ of a graph $G$, transforms $\alpha$ into $\beta$ by recolouring one vertex at a time, such that after each recolouring step we again have a proper $k$-colouring of $G$. The…
Quality gain is the expected relative improvement of the function value in a single step of a search algorithm. Quality gain analysis reveals the dependencies of the quality gain on the parameters of a search algorithm, based on which one…
I describe an analytical approximation for calculating the short-term probability of loss of a chromosome under the neutral Wright-Fisher model with recombination. I also present an upper and lower bound for this probability. Exact…
In condensed matter physics, simplified descriptions are obtained by coarse-graining the features of a system at a certain characteristic length, defined as the typical length beyond which some properties are no longer correlated. From a…
The Human Genome Project (HGP) provides researchers with the data of nearly all human genes and the challenge to use this information for elucidating the etiology of common disorders. A secondary Darwinian method was applied to HGP and…
Understanding how molecular changes caused by genetic variation drive disease risk is crucial for deciphering disease mechanisms. However, interpreting genome sequences is challenging because of the vast size of the human genome, and…
We investigate joint spectral characteristics of a family of matrices $\mathcal F $, associated with products in the semigroup generated by $\mathcal F$. In the literature, extremal measures such as the well-known joint spectral radius and…
Being able to store and transmit human genome sequences is an important part in genomic research and industrial applications. The complete human genome has 3.1 billion base pairs (haploid), and storing the entire genome naively takes about…
DNA renaturation is the recombination of two complementary single strands to form a double helix. It is experimentally known that renaturation proceeds through the formation of a double stranded nucleus of several base pairs (the rate…
Convolutions of independent random variables often arise in a natural way in many applied problems. In this article, we compare convolutions of two sets of gamma (negative binomial) random variables in the convolution order and the usual…
This paper investigates the influence of genotype size on evolutionary algorithms' performance. We consider genotype compression (where genotype is smaller than phenotype) and expansion (genotype is larger than phenotype) and define…
Biological cells replicate their genomes in a well-planned manner. The DNA replication program of an organism determines the timing at which different genomic regions are replicated, with fundamental consequences for cell homeostasis and…
Multi-trait genome-wide association studies (GWAS) use multi-variate statistical methods to identify associations between genetic variants and multiple correlated traits simultaneously, and have higher statistical power than independent…
The most common gene regulation mechanism is when a transcription factor protein binds to a regulatory sequence to increase or decrease RNA transcription. However, transcription factors face two main challenges when searching for these…
One of the key difficulties in using estimation-of-distribution algorithms is choosing the population size(s) appropriately: Too small values lead to genetic drift, which can cause enormous difficulties. In the regime with no genetic drift,…
The lengths of the telomere regions of chromosomes in a population of cells are modelled using a chemical master equation formalism, from which the evolution of the average number of cells of each telomere length is extracted. In…