中文
相关论文

相关论文: Simplified amino acid alphabets based on deviation…

200 篇论文

A new perspective is introduced regarding the analysis of Multiple Sequence Alignments (MSA), representing aligned data defined over a finite alphabet of symbols. The framework is designed to produce a block decomposition of an MSA, where…

信息论 · 计算机科学 2024-01-11 Christopher Barrett , Andrei Bura , Fenix Huang , Christian Reidys

The study is aimed at revealing the most important substructures (fragments) of polyenes with heteroatoms determining the alteration in the conjugation energy of the whole compound due to substitution and the relevant charge redistribution.…

化学物理 · 物理学 2022-02-16 Viktorija Gineityte

The similarity between protein sequences is a directly and easly computed quantity from which to deduce information about their evolutionary distance and to detect homologous proteins. The SIMAP database -- Similarity Matrix of Proteins --…

定量方法 · 定量生物学 2007-05-23 Concetta Miccio , Thomas Rattei

This paper presents results of a study of the performance of several base classifiers for recognition of handwritten characters of the modern Latin alphabet. Base classification performance is further enhanced by utilizing Viterbi error…

计算机视觉与模式识别 · 计算机科学 2021-11-30 Hélder Campos , Nuno Paulino

The paper represents three supplements to the source paper, q-bio/0610044 [q-bio.OT], with three new series of harmonic structures of the genetic code, determined by Gauss arithmetical algorithm; by Table of Minimal Adding, as in…

其他定量生物学 · 定量生物学 2018-02-16 Miloje M. Rakocevic

Decoding sequences that stem from multiple transmissions of a codeword over an insertion, deletion, and substitution channel is a critical component of efficient deoxyribonucleic acid (DNA) data storage systems. In this paper, we consider a…

System identification is normally involved in augmenting time series data by time shifting and nonlinearisation (e.g., polynomial basis), both of which introduce redundancy in features and samples. Many research works focus on reducing…

机器学习 · 计算机科学 2025-09-05 Tingna Wang , Sikai Zhang , Mingming Song , Limin Sun

Nuclear Magnetic Resonance (NMR) is used in structural biology to experimentally determine the structure of proteins, which is used in many areas of biology and is an important part of drug development. Unfortunately, NMR data can cost…

定量方法 · 定量生物学 2022-08-03 Jia Qi Yip , Dianwen Ng , Bin Ma , Konstantin Pervushin , Eng Siong Chng

The problem of determining the correct order of fluctuation of the optimal alignment score of two random strings of length $n$ has been open for several decades. It is known that the biased expected effect of a random letter-change on the…

概率论 · 数学 2012-11-26 Saba Amsalu , Raphael Hauser , Heinrich Matzinger

Biomedical research papers use significantly different language and jargon when compared to typical English text, which reduces the utility of pre-trained NLP models in this domain. Meanwhile Medline, a database of biomedical abstracts,…

机器学习 · 计算机科学 2021-09-15 Justin Sybrandt , Ilya Safro

Upon the covalent-bonding hybrid of the nitrogen atoms taken as a measure for the structural regularity in nucleobases, it can be identified that the internal relation within the 20 amino acids follows a cooperative vector-in-space addition…

生物大分子 · 定量生物学 2007-05-23 Chi Ming Yang

The identification of low-energy conformers for a given molecule is a fundamental problem in computational chemistry and cheminformatics. We assess here a conformer search that employs a genetic algorithm for sampling the low-energy segment…

生物大分子 · 定量生物学 2015-11-24 Adriana Supady , Volker Blum , Carsten Baldauf

The Poisson-sampling technique eliminates dependencies among symbol appearances in a random sequence. It has been used to simplify the analysis and strengthen the performance guarantees of randomized algorithms. Applying this method to…

信息论 · 计算机科学 2014-05-30 Jayadev Acharya , Ashkan Jafarpour , Alon Orlitsky , Ananda Theertha Suresh

For many interesting tasks, such as medical diagnosis and web page classification, a learner only has access to some positively labeled examples and many unlabeled examples. Learning from this type of data requires making assumptions about…

机器学习 · 计算机科学 2018-08-28 Jessa Bekker , Jesse Davis

We present methods for calculating a measure of phonotactic complexity---bits per phoneme---that permits a straightforward cross-linguistic comparison. When given a word, represented as a sequence of phonemic segments such as symbols in the…

计算与语言 · 计算机科学 2020-05-11 Tiago Pimentel , Brian Roark , Ryan Cotterell

We have investigated the deviation from the standard recombination process, using the ACBAR 2008 and the WMAP 3 year data. In this investigation, we have considered the possibility of the accelerated recombination as well as the delayed…

天体物理学 · 物理学 2009-11-13 Jaiseung Kim , Pavel Naselsky

We demonstrate a new algorithm for finding protein conformations that minimize a non-bonded energy function. The new algorithm, called the difference map, seeks to find an atomic configuration that is simultaneously in two constraint…

生物大分子 · 定量生物学 2007-06-13 Ivan C. Rankenburg , Veit Elser

The huge amount of data acquired by high-throughput sequencing requires data reduction for effective analysis. Here we give a clustering algorithm for genome-wide open chromatin data using a new data reduction method. This method regards…

Bloom filter is a compact memory-efficient probabilistic data structure supporting membership testing, i.e., to check whether an element is in a given set. However, as Bloom filter maps each element with uniformly random hash functions, few…

数据库 · 计算机科学 2021-06-15 Rongbiao Xie , Meng Li , Zheyu Miao , Rong Gu , He Huang , Haipeng Dai , Guihai Chen

Motivated by the randomized generation of slowly synchronizing automata, we study automata made of permutation letters and a merging letter of rank $ n\!-\!1 $. We present a constructive randomized procedure to generate synchronizing…

形式语言与自动机理论 · 计算机科学 2018-06-27 Costanza Catalano , Raphaël M. Jungers