English
Related papers

Related papers: Identification of Protein Coding Regions in Genomi…

200 papers

Algorithms that detect covariance between pairs of columns in multiple sequence alignments are commonly employed to predict functionally important residues and structural contacts. However, the assumption that co-variance only occurs…

Quantitative Methods · Quantitative Biology 2014-01-07 Kyle E. Kreth , Anthony A. Fodor

Automatic detecting anomalous regions in images of objects or textures without priors of the anomalies is challenging, especially when the anomalies appear in very small areas of the images, making difficult-to-detect visual variations,…

Computer Vision and Pattern Recognition · Computer Science 2022-07-05 Jie Yang , Yong Shi , Zhiquan Qi

DNA exhibits remarkable potential as a data storage solution due to its impressive storage density and long-term stability, stemming from its inherent biomolecular structure. However, developing this novel medium comes with its own set of…

Image and Video Processing · Electrical Eng. & Systems 2023-09-14 Trung Hieu Le , Xavier Pic , Jeremy Mateos , Marc Antonini

Recently several deep learning models have been used for DNA sequence based classification tasks. Often such tasks require long and variable length DNA sequences in the input. In this work, we use a sequence-to-sequence autoencoder model to…

Genomics · Quantitative Biology 2019-06-10 Vishal Agarwal , N Jayanth Kumar Reddy , Ashish Anand

Applying the knowledge of an object detector trained on a specific domain directly onto a new domain is risky, as the gap between two domains can severely degrade model's performance. Furthermore, since different instances commonly embody…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Minghao Xu , Hang Wang , Bingbing Ni , Qi Tian , Wenjun Zhang

Computational methods for discovering patterns of local correlations in sequences are important in computational biology. Here we show how to determine the optimal partitioning of aligned sequences into non-overlapping segments such that…

Computational Engineering, Finance, and Science · Computer Science 2012-06-26 Joseph Bockhorst , Nebojsa Jojic

Through sequence-based classification, this paper tries to accurately predict the DNA binding sites of transcription factors (TFs) in an unannotated cellular context. Related methods in the literature fail to perform such predictions…

Machine Learning · Computer Science 2016-11-17 Ritambhara Singh , Jack Lanchantin , Gabriel Robins , Yanjun Qi

This paper designs an efficient two-class pattern classifier utilizing asynchronous cellular automata (ACAs). The two-state three-neighborhood one-dimensional ACAs that converge to fixed points from arbitrary seeds are used here for pattern…

Cellular Automata and Lattice Gases · Physics 2016-02-01 Biswanath Sethi , Souvik Roy , Sukanta Das

Molecular data from tumor profiles is high dimensional. Tumor profiles can be characterized by tens of thousands of gene expression features. Due to the size of the gene expression feature set machine learning methods are exposed to noisy…

Machine Learning · Computer Science 2020-07-14 Martin Palazzo , Pierre Beauseroy , Patricio Yankilevich

We introduce a novel method to analyse complete genomes and recognise some distinctive features by means of an adaptive compression algorithm, which is not DNA-oriented. We study the Information Content as a function of the number of…

Genomics · Quantitative Biology 2007-05-23 Giulia Menconi

Motivation: With the development of third-generation sequencing technologies, people are able to obtain DNA sequences with lengths from 10s to 100s of kb. These long reads allow protein domain annotation without assembly, thus can produce…

Genomics · Quantitative Biology 2021-07-09 Du Nan , Jiayu Shang , Yanni Sun

Biomolecular computation has emerged as an important area of computer science research due to its high information density, immense parallelism opportunity along with potential applications in cryptography, genetic engineering and…

Formal Languages and Automata Theory · Computer Science 2022-06-07 Anupam Chattopadhyay , Arnab Chakrabarti

The aim of this study was to develop a method that would identify the cluster centroids and the optimal number of clusters for a given sensitivity level and could work equally well for the different sequence datasets. A novel method that…

Genomics · Quantitative Biology 2023-12-01 Manal Helal , Fanrong Kong , Sharon C-A Chen , Fei Zhou , Dominic E Dwyer , John Potter , Vitali Sintchenko

Proteins populate a manifold in the high-dimensional sequence space whose geometrical structure guides their natural evolution. Leveraging recently-developed structure prediction tools based on transformer models, we first examine the…

Biomolecules · Quantitative Biology 2023-11-13 A. Zambon , R. Zecchina , G. Tiana

Lineage tracing, the determination and mapping of progeny arising from single cells, is an important approach enabling the elucidation of mechanisms underlying diverse biological processes ranging from development to disease. We developed a…

In this paper we study the structure of specific linear codes called DNA codes. The first attempts on studying such codes have been proposed over four element rings which are naturally matched with DNA four letters. Later, double (pair) DNA…

Combinatorics · Mathematics 2017-03-31 Fatmanur Gürsoy , Elif Segah Oztas , Irfan Siap

Automatic code generation for low-dimensional geometric algorithms is capable of producing efficient low-level software code through a high-level geometric domain specific language. Geometric Algebra (GA) is one of the most suitable…

Mathematical Software · Computer Science 2016-07-19 Ahmad Hosney Awad Eid

Capturing the physical organisation and dynamics of genomic regions is one of the major open challenges in biology. The kinetoplast DNA (kDNA) is a topologically complex genome, made by thousands of DNA (mini and maxi) circles interlinked…

Soft Condensed Matter · Physics 2025-04-16 Saminathan Ramakrishnan , Auro Varat Patnaik , Guglielmo Grillo , Luca Tubiana , Davide Michieletto

Principal Components Analysis (PCA) is a common way to study the sources of variation in a high-dimensional data set. Typically, the leading principal components are used to understand the variation in the data or to reduce the dimension of…

Most approaches to prediction of protein function from primary structure are based on similarity between the query sequence and sequences of known function. This approach, however, disregards the occurrence of gene duplication (paralogy) or…

Populations and Evolution · Quantitative Biology 2014-04-03 Paulo Bandiera-Paiva , Jackson C. Lima , Marcelo R. S. Briones