English
Related papers

Related papers: Information profiles for DNA pattern discovery

200 papers

Long-context DNA models are limited by token-mixing cost and by how compression allocates representational budget across the genome. Existing approaches operate close to base-pair resolution, apply fixed downsampling, or learn…

Genomics · Quantitative Biology 2026-05-13 Jianan Zhao , Xixian Liu , Zhihao Zhan , Xinyu Yuan , Hongyu Guo , Jian Tang

Finite-context models (FCMs) are widely used for compressing symbolic sequences such as DNA, where predictive performance depends critically on the context length k and smoothing parameter {\alpha}. In practice, these hyperparameters are…

Machine Learning · Statistics 2026-03-23 José Contente , Ana Martins , Armando J. Pinho , Sónia Gouveia

Motivation Protein fold recognition is an important problem in structural bioinformatics. Almost all traditional fold recognition methods use sequence (homology) comparison to indirectly predict the fold of a tar get protein based on the…

Machine Learning · Computer Science 2017-06-06 Jie Hou , Badri Adhikari , Jianlin Cheng

Numerous temporal inference tasks such as fault monitoring and anomaly detection exhibit a persistence property: for example, if something breaks, it stays broken until an intervention. When modeled as a Dynamic Bayesian Network,…

Artificial Intelligence · Computer Science 2012-06-18 Tomas Singliar , Denver Dash

Motif finding is an important step for the detection of rare events occurring in a set of DNA or protein sequences. Extraction of information about these rare events can lead to new biological discoveries. Motifs are some important patterns…

Data Structures and Algorithms · Computer Science 2024-03-04 Saurav Dhar , Amlan Saha , Dhiman Goswami , Md. Abul Kashem Mia

The interaction between proteins and DNA is a key driving force in a significant number of biological processes such as transcriptional regulation, repair, recombination, splicing, and DNA modification. The identification of DNA-binding…

Quantitative Methods · Quantitative Biology 2017-05-10 Hamid Reza Hassanzadeh , Pushkar Kolhe , Charles L. Isbell , May D. Wang

Gradual pattern mining allows for extraction of attribute correlations through gradual rules such as: "the more X, the more Y". Such correlations are useful in identifying and isolating relationships among the attributes that may not be…

Databases · Computer Science 2021-06-29 Dickson Odhiambo Owuor

To simulate long time and length scale processes involving DNA it is necessary to use a coarse-grained description. Here we provide an overview of different approaches to such coarse graining, focussing on those at the nucleotide level that…

The identification of repeating patterns in discrete grids is rudimentary within symbolic reasoning, algorithm synthesis and structural optimization across diverse computational domains. Although statistical approaches targeting noisy data…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Sushish Baral , Paulo Garcia , Warisa Sritriratanarak

In this paper we describe a new technique for the comparison of populations of DNA strands. Comparison is vital to the study of ecological systems, at both the micro and macro scales. Existing methods make use of DNA sequencing and cloning,…

Biomolecules · Quantitative Biology 2008-07-02 Dennis Shasha , Martyn Amos

Identifiability is a necessary condition for successful parameter estimation of dynamic system models. A major component of identifiability analysis is determining the identifiable parameter combinations, the functional forms for the…

Quantitative Methods · Quantitative Biology 2013-10-07 Marisa C. Eisenberg , Michael A. L. Hayashi

DNA-based storage offers unprecedented density and durability, but its scalability is fundamentally limited by the efficiency of parallel strand synthesis. Existing methods either allow unconstrained nucleotide additions to individual…

Information Theory · Computer Science 2025-10-27 Boaz Moav , Ryan Gabrys , Eitan Yaakobi

When analyzing communities of microorganisms from their sequenced DNA, an important task is taxonomic profiling: enumerating the presence and relative abundance of all organisms, or merely of all taxa, contained in the sample. This task can…

Genomics · Quantitative Biology 2020-01-24 Simon Foucart , David Koslicki

Solving semiparametric models can be computationally challenging because the dimension of parameter space may grow large with increasing sample size. Classical Newton's method becomes quite slow and unstable with intensive calculation of…

Computation · Statistics 2021-08-19 Yucong Lin , Jinhua Su , Yang Liu , Jue Hou , Feifei Wang

This paper presents tailor-made neural model structures and two custom fitting criteria for learning dynamical systems. The proposed framework is based on a representation of the system behavior in terms of continuous-time state-space…

Systems and Control · Electrical Eng. & Systems 2021-09-02 Marco Forgione , Dario Piga

In microarray experiments, it is often of interest to identify genes which have a pre-specified gene expression profile with respect to time. Methods available in the literature are, however, typically not stringent enough in identifying…

Applications · Statistics 2009-01-18 J. Tuke , G. F. V. Glonek , P. J. Solomon

We study the problem of image registration in the finite-resolution regime and characterize the error probability of algorithms as a function of properties of the transformation and the image capture noise. Specifically, we define a…

Information Theory · Computer Science 2020-01-14 Ravi Kiran Raman , Lav R. Varshney

This paper describes two approaches for content-based image retrieval and pattern spotting in document images using deep learning. The first approach uses a pre-trained CNN model to cope with the lack of training data, which is fine-tuned…

Computer Vision and Pattern Recognition · Computer Science 2019-07-25 Kelly Lais Wiggers , Alceu de Souza Britto Junior , Alessandro Lameiras Koerich , Laurent Heutte , Luiz Eduardo Soares de Oliveira

Biclustering is an unsupervised data mining technique that aims to unveil patterns (biclusters) from gene expression data matrices. In the framework of this thesis, we propose new biclustering algorithms for microarray data. The latter is…

Machine Learning · Computer Science 2018-11-26 Amina Houari

Finite mixture models are a useful statistical model class for clustering and density approximation. In the Bayesian framework finite mixture models require the specification of suitable priors in addition to the data model. These priors…

Methodology · Statistics 2024-07-09 Bettina Grün , Gertraud Malsiner-Walli