中文
相关论文

相关论文: Sequential Labelling and DNABERT For Splice Site P…

200 篇论文

Splice sites play a crucial role in gene expression, and accurate prediction of these sites in DNA sequences is essential for diagnosing and treating genetic disorders. We address the challenge of splice site prediction by introducing…

基因组学 · 定量生物学 2023-11-23 Asmita Poddar , Vladimir Uzun , Elizabeth Tunbridge , Wilfried Haerty , Alejo Nevado-Holgado

The primary step in search of the gene prediction is an identification of the coding region from genomic DNA sequence. Gene structure in the case of a eukaryotic organism is composed of promoter, intron, start codon, exons, stop codon, etc.…

基因组学 · 定量生物学 2019-07-23 Srabanti Maji , Soumen Kanrar

A eukaryotic gene consists of multiple exons (protein coding regions) and introns (non-coding regions), and a splice junction refers to the boundary between a pair of exon and intron. Precise identification of spice junctions on a gene is…

机器学习 · 计算机科学 2015-12-17 Byunghan Lee , Taehoon Lee , Byunggook Na , Sungroh Yoon

Background: Exonic splice enhancers are sequences embedded within exons which promote and regulate the splicing of the transcript in which they are located. A class of exonic splice enhancers are the SR proteins, which are thought to…

基因组学 · 定量生物学 2007-05-23 Thomas A. Down , Bernard Leong , Tim J. P. Hubbard

Bioinformatics encompass storing, analyzing and interpreting the biological data. Most of the challenges for Machine Learning methods like Cellular Automata is to furnish the functional information with the corresponding biological…

计算工程、金融与科学 · 计算机科学 2014-04-25 Pokkuluri Kiran Sree , Inampudi Ramesh Babu , SSSN Usha Devi N

Recently several deep learning models have been used for DNA sequence based classification tasks. Often such tasks require long and variable length DNA sequences in the input. In this work, we use a sequence-to-sequence autoencoder model to…

基因组学 · 定量生物学 2019-06-10 Vishal Agarwal , N Jayanth Kumar Reddy , Ashish Anand

Motivation: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alignment demands…

基因组学 · 定量生物学 2025-09-23 Siying Yang , Neng Huang , Heng Li

Several modern genomic technologies, such as DNA-Methylation arrays, measure spatially registered probes that number in the hundreds of thousands across multiplechromosomes. The measured probes are by themselves less interesting…

应用统计 · 统计学 2016-11-16 John Nagorski , Genevera I. Allen

We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e., DNA…

基因组学 · 定量生物学 2024-10-23 Zhihan Zhou , Weimin Wu , Harrison Ho , Jiayi Wang , Lizhen Shi , Ramana V Davuluri , Zhong Wang , Han Liu

Sequence labeling (SL) is a fundamental research problem encompassing a variety of tasks, e.g., part-of-speech (POS) tagging, named entity recognition (NER), text chunking, etc. Though prevalent and effective in many downstream applications…

计算与语言 · 计算机科学 2020-11-16 Zhiyong He , Zanbo Wang , Wei Wei , Shanshan Feng , Xianling Mao , Sheng Jiang

The problem of identifying splice sites consists of two sub-problems: finding their boundaries, and characterizing their sequence markers. Other splicing elements---including, enhancers and silencers---that occur in the intronic and exonic…

定量方法 · 定量生物学 2012-06-27 Christine Lo , Boyko Kakaradov , Daniel Lokshtanov , Christina Boucher

DNA sequencing to identify genetic variants is becoming increasingly valuable in clinical settings. Assessment of variants in such sequencing data is commonly implemented through Bayesian heuristic algorithms. Machine learning has shown…

Gene annotation has traditionally required direct comparison of DNA sequences between an unknown gene and a database of known ones using string comparison methods. However, these methods do not provide useful information when a gene does…

机器学习 · 计算机科学 2019-09-17 James K. Senter , Taylor M. Royalty , Andrew D. Steen , Amir Sadovnik

Labeling of DNA molecules is a fundamental technique for DNA visualization and analysis. This process was mathematically modeled in [1], where the received sequence indicates the positions of the used labels. In this work, we develop error…

信息论 · 计算机科学 2025-11-04 Dganit Hanania , Eitan Yaakobi

We develop statistically based methods to detect single nucleotide DNA mutations in next generation sequencing data. Sequencing generates counts of the number of times each base was observed at hundreds of thousands to billions of genome…

应用统计 · 统计学 2012-10-01 Omkar Muralidharan , Georges Natsoulis , John Bell , Hanlee Ji , Nancy R. Zhang

Bioinformatics, as an emerging and rapidly developing interdisciplinary, has become a promising and popular research field in 21st century. Extracting and explaining useful biological information from huge amount of genetic data is an…

应用统计 · 统计学 2014-11-25 Beilin Jia , Wenli Shi , Feng Zhang

In ecology it has become common to apply DNA barcoding to biological samples leading to datasets containing a large number of nucleotide sequences. The focus is then on inferring the taxonomic placement of each of these sequences by…

应用统计 · 统计学 2022-01-25 Alessandro Zito , Tommaso Rigon , David B. Dunson

Identifying gene splicing is a core and significant task confronted in modern collaboration between artificial intelligence and bioinformatics. Past decades have witnessed great efforts on this concern, such as the bio-plausible splicing…

定量方法 · 定量生物学 2024-06-19 Qi-Jie Li , Qian Sun , Shao-Qun Zhang

Named entity recognition (NER) is a widely studied task in natural language processing. Recently, a growing number of studies have focused on the nested NER. The span-based methods, considering the entity recognition as a span…

计算与语言 · 计算机科学 2021-06-22 Zeqi Tan , Yongliang Shen , Shuai Zhang , Weiming Lu , Yueting Zhuang

Circular RNAs (circRNAs) are important components of the non-coding RNA regulatory network. Previous circRNA identification primarily relies on high-throughput RNA sequencing (RNA-seq) data combined with alignment-based algorithms that…

机器学习 · 计算机科学 2025-07-14 Tianyou Jiang
‹ 上一页 1 2 3 10 下一页 ›