English
Related papers

Related papers: Geometric approach to string analysis: deviation f…

200 papers

Recent advances in high-throughput genomics technologies have resulted in the sequencing of large numbers of (near) complete genomes. These genome sequences are being mined for important functional elements, such as genes. They are also…

Genomics · Quantitative Biology 2007-05-23 Lior Pachter

Grouping elements into families to analyse them separately is a standard analysis procedure in many areas of sciences. We propose herein a new algorithm based on the simple idea that members from a family look like each other, and don't…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Axel Descamps , Sélène Forget , Aliénor Lahlou , Claire Lavergne , Camille Berthelot , Guillaume Stirnemann , Rodolphe Vuilleumier , Nicolas Chéron

How a single fertilized cell gives rise to a complex array of specialized cell types in development is a central question in biology. The cells grow, divide, and acquire differentiated characteristics through poorly understood molecular…

Machine Learning · Computer Science 2025-03-26 Da Kuang , Guanwen Qiu , Junhyong Kim

Probabilistic models over strings have played a key role in developing methods allowing indels to be treated as phylogenetically informative events. There is an extensive literature on using automata and transducers on phylogenies to do…

Populations and Evolution · Quantitative Biology 2013-07-15 Alexandre Bouchard-Côté

Complex interactions between genes or proteins contribute a substantial part to phenotypic evolution. Here we develop an evolutionarily grounded method for the cross-species analysis of interaction networks by {\em alignment}, which maps…

Molecular Networks · Quantitative Biology 2009-11-13 Johannes Berg , Michael Lässig

Quantitative methods for studying biodiversity have been traditionally rooted in the classical theory of finite frequency tables analysis. However, with the help of modern experimental tools, like high throughput sequencing, we now begin to…

Methodology · Statistics 2015-12-22 Maciej Pietrzak , Grzegorz A. Rempała , Michał Seweryn , Jacek Wesołowski

The normalized edit distance is one of the distances derived from the edit distance. It is useful in some applications because it takes into account the lengths of the two strings compared. The normalized edit distance is not defined in…

Neural and Evolutionary Computing · Computer Science 2013-12-09 Muhammad Marwan Muhammad Fuad

We propose a new alignment-free algorithm by constructing a compact vector representation on $\mathbb{R}^{24}$ of a DNA sequence of arbitrary length. Each component of this vector is obtained from a representative sequence, the elements of…

Data Structures and Algorithms · Computer Science 2024-09-27 Probir Mondal , Pratyay Banerjee , Debranjan Pal , Krishnendu Basuli

RNA sequencing (RNA-seq) enables characterization and quantification of individual transcriptomes as well as detection of patterns of allelic expression and alternative splicing. Current RNA-seq protocols depend on high-throughput…

Genomics · Quantitative Biology 2015-06-19 Hyunghoon Cho , Joe Davis , Xin Li , Kevin S. Smith , Alexis Battle , Stephen B. Montgomery

The prediction of phenotypic traits using high-density genomic data has many applications such as the selection of plants and animals of commercial interest; and it is expected to play an increasing role in medical diagnostics. Statistical…

Methodology · Statistics 2016-09-29 Marco Scutari , Ian Mackay , David Balding

In this paper, we present an optical computing method for string data alignment applicable to genome information analysis. By applying moire technique to spatial encoding patterns of deoxyribonucleic acid (DNA) sequences, association…

Gene expression estimation from pathology images has the potential to reduce the RNA sequencing cost. Point-wise loss functions have been widely used to minimize the discrepancy between predicted and absolute gene expression values.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Kazuya Nishimura , Haruka Hirose , Ryoma Bise , Kaito Shiku , Yasuhiro Kojima

A novel non-linear approach to fast and effective comparison of sequences is presented, compared to the traditional cross-correlation operator, and illustrated with respect to DNA sequences.

Computer Vision and Pattern Recognition · Computer Science 2007-05-23 Luciano da Fontoura Costa

We consider the problem of estimating the evolutionary history of a set of species (phylogeny or species tree) from several genes. It is known that the evolutionary history of individual genes (gene trees) might be topologically distinct…

Populations and Evolution · Quantitative Biology 2016-11-18 Gautam Dasarathy , Robert Nowak , Sebastien Roch

We present a general setting for structure-sequence comparison in a large class of RNA structures that unifies and generalizes a number of recent works on specific families on structures. Our approach is based on tree decomposition of…

Quantitative Methods · Quantitative Biology 2012-06-21 Philippe Rinaudo , Yann Ponty , Dominique Barth , Alain Denise

We analyse a simple discrete-time stochastic process for the theoretical modeling of the evolution of protein lengths. At every step of the process a new protein is produced as a modification of one of the proteins already existing and its…

Populations and Evolution · Quantitative Biology 2009-11-13 C. Destri , C. Miccio

We perform differential expression analysis of high-throughput sequencing count data under a Bayesian nonparametric framework, removing sophisticated ad-hoc pre-processing steps commonly required in existing algorithms. We propose to use…

Applications · Statistics 2017-05-04 Siamak Zamani Dadaneh , Xiaoning Qian , Mingyuan Zhou

Motivation: Most existing methods for DNA sequence analysis rely on accurate sequences or genotypes. However, in applications of the next-generation sequencing (NGS), accurate genotypes may not be easily obtained (e.g. multi-sample…

Genomics · Quantitative Biology 2013-03-19 Heng Li

Real bipartite networks combine degree-constrained random mixing with structured, locality-like rules. We introduce a statistical filter that benchmarks node-level bipartite clustering against degree-preserving randomizations to classify…

Physics and Society · Physics 2025-09-27 Lucía S. Ramírez , Roya Aliakbarisani , M. Ángeles Serrano , Marián Boguñá

We consider the problem of coding for the substring channel, in which information strings are observed only through their (multisets of) substrings. Due to existing DNA sequencing techniques and applications in DNA-based storage systems,…

Information Theory · Computer Science 2024-03-27 Yonatan Yehezkeally , Nikita Polyanskii