English
Related papers

Related papers: Optimizing Smith-Waterman alignments

200 papers

Motivation: Standard algorithms for pairwise protein sequence alignment make the simplifying assumption that amino acid substitutions at neighboring sites are uncorrelated. This assumption allows implementation of fast algorithms for…

Biomolecules · Quantitative Biology 2014-05-27 Gavin E. Crooks , Richard E. Green , Steven E. Brenner

We propose a new segmentation evaluation metric, called segmentation similarity (S), that quantifies the similarity between two segmentations as the proportion of boundaries that are not transformed when comparing them using edit distance,…

Computation and Language · Computer Science 2012-06-08 Chris Fournier , Diana Inkpen

In this report, correlation of the pixels comprising a microarray spot is investigated. Subsequently, correlation statistics namely: Pearson correlation and Spearman rank correlation are used to segment the foreground and background…

Genomics · Quantitative Biology 2007-05-23 Radhakrishnan Nagarajan , Meenakshi Upreti

Traditional pairwise sequence alignment is based on matching individual samples from two sequences, under time monotonicity constraints. However, in many application settings matching subsequences (segments) instead of individual samples…

Databases · Computer Science 2016-09-28 Shahriar Shariat , Vladimir Pavlovic

Recent advances in unsupervised domain adaptation mainly focus on learning shared representations by global distribution alignment without considering class information across domains. The neglect of class information, however, may lead to…

Computer Vision and Pattern Recognition · Computer Science 2019-04-16 Chao Chen , Zhihang Fu , Zhihong Chen , Zhaowei Cheng , Xinyu Jin , Xian-Sheng Hua

We propose local segmentation of multiple sequences sharing a common time- or location-index, building upon the single sequence local segmentation methods of Niu and Zhang (2012) and Fang, Li and Siegmund (2016). We also propose reverse…

Statistics Theory · Mathematics 2022-01-10 Hock-Peng Chan , Hao Chen

Considering two optimally aligned random sequences, we investigate the effect on the alignment score caused by changing a random letter in one of the two sequences. Using this idea in conjunction with large deviations theory, we show that…

Probability · Mathematics 2015-06-15 Raphael Hauser , Heinrich Matzinger

This paper describes a new alignment algorithm for sequences that can be used for determination of deletions and substitutions. It provides several solutions out of which the best one can be chosen on the basis of minimization of gaps or…

Information Theory · Computer Science 2012-11-01 Sandeep Hosangadi , Subhash Kak

We consider DNA codes based on the nearest-neighbor (stem) similarity model which adequately reflects the "hybridization potential" of two DNA sequences. Our aim is to present a survey of bounds on the rate of DNA codes with respect to a…

Information Theory · Computer Science 2016-11-17 A. D'yachkov , A. Voronina , A. Macula , T. Renz , V. Rykov

The subject of this paper is to study conformance checking for timed models, that is, process models that consider both the sequence of events in a process as well as the timestamps at which each event is recorded. Time-aware process mining…

Formal Languages and Automata Theory · Computer Science 2022-07-06 Thomas Chatain , Neha Rino

We present a sequence-based probabilistic formalism that directly addresses co-operative effects in networks of interacting positions in proteins, providing significantly improved contact prediction, as well as accurate quantitative…

Quantitative Methods · Quantitative Biology 2012-07-12 Alan Lapedes , Bertrand Giraud , Christopher Jarzynski

Gene annotation has traditionally required direct comparison of DNA sequences between an unknown gene and a database of known ones using string comparison methods. However, these methods do not provide useful information when a gene does…

Machine Learning · Computer Science 2019-09-17 James K. Senter , Taylor M. Royalty , Andrew D. Steen , Amir Sadovnik

A new method to simulate probability distributions in regions where the events are VERY unlikely (e.g. p ~ 10^{-40}) is presented. The basic idea is to represent the underlying probability space by the phase space of a physical system. The…

Disordered Systems and Neural Networks · Physics 2009-11-07 Alexander K. Hartmann

A profile comparison method with position-specific scoring matrix (PSSM) is one of the most accurate alignment methods. Currently, cosine similarity and correlation coefficient are used as scoring functions of dynamic programming to…

Quantitative Methods · Quantitative Biology 2018-02-20 Kazunori D Yamada

Score matching provides an effective approach to learning flexible unnormalized models, but its scalability is limited by the need to evaluate a second-order derivative. In this paper, we present a scalable approximation to a general family…

Machine Learning · Statistics 2020-02-19 Ziyu Wang , Shuyu Cheng , Yueru Li , Jun Zhu , Bo Zhang

The subject of this paper is to study conformance checking for timed models, that is, process models that consider both the sequence of events in a process as well as the timestamps at which each event is recorded. Time-aware process mining…

Formal Languages and Automata Theory · Computer Science 2022-10-28 Neha Rino , Thomas Chatain

Summary: The Smith Waterman (SW) algorithm, which produces the optimal pairwise alignment between two sequences, is frequently used as a key component of fast heuristic read mapping and variation detection tools, but current implementations…

Genomics · Quantitative Biology 2014-03-05 Mengyao Zhao , Wan-Ping Lee , Erik Garrison , Gabor T. Marth

We explore connections between metagenomic read assignment and the quantification of transcripts from RNA-Seq data. In particular, we show that the recent idea of pseudoalignment introduced in the RNA-Seq context is suitable in the…

Quantitative Methods · Quantitative Biology 2015-12-02 Lorian Schaeffer , Harold Pimentel , Nicolas Bray , Páll Melsted , Lior Pachter

One of the most common methods for statistical inference is the maximum likelihood estimator (MLE). The MLE needs to compute the normalization constant in statistical models, and it is often intractable. Using unnormalized statistical…

Statistics Theory · Mathematics 2016-04-26 Takafumi Kanamori , Takashi Takenouchi

We study the problem of local alignment, which is finding pairs of similar subsequences with gaps. The problem exists in biosequence databases. BLAST is a typical software for finding local alignment based on heuristic, but could miss…

Databases · Computer Science 2012-08-02 Xiaochun Yang , Honglei Liu , Bin Wang