中文
相关论文

相关论文: HyenaDNA: Long-Range Genomic Sequence Modeling at …

200 篇论文

Pre-training large language models on genomic sequences is a powerful approach for learning biologically meaningful representations. Masked language modeling (MLM) methods, such as DNABERT and Nucleotide Transformer (NT), achieve strong…

基因组学 · 定量生物学 2025-08-20 Ke Ding , Brian Parker , Jiayu Wen

Each human genome is a 3 billion base pair set of encoding instructions. Decoding the genome using deep learning fundamentally differs from most tasks, as we do not know the full structure of the data and therefore cannot design…

机器学习 · 计算机科学 2016-05-24 Laura Deming , Sasha Targ , Nate Sauder , Diogo Almeida , Chun Jimmie Ye

Several processes in the cell, such as gene regulation, start when key proteins recognise and bind to short DNA sequences. However, as these sequences can be hundreds of million times shorter than the genome, they are hard to find by simple…

亚细胞过程 · 定量生物学 2021-01-27 Markus Nyberg , Tobias Ambjörnsson , Per Stenberg , and Ludvig Lizana

Background: Small interfering RNA (siRNA) is a promising therapeutic agent due to its ability to silence disease-related genes via RNA interference. While traditional machine learning and early deep learning methods have made progress in…

生物大分子 · 定量生物学 2025-03-07 Wangdan Liao , Weidong Wang

Recent advances in attention-free sequence models rely on convolutions as alternatives to the attention operator at the core of Transformers. In particular, long convolution sequence models have achieved state-of-the-art performance in many…

Many machine learning models use the manipulation of dimensions as a driving force to enable models to identify and learn important features in data. In the case of sequential data this manipulation usually happens on the token dimension…

机器学习 · 计算机科学 2023-10-24 Daniel Biermann , Fabrizio Palumbo , Morten Goodwin , Ole-Christoffer Granmo

Accurate phenotype prediction from RNA sequencing (RNA-seq) data is essential for diagnosis, biomarker discovery, and personalized medicine. Deep learning models have demonstrated strong potential to outperform classical machine learning…

机器学习 · 计算机科学 2026-03-10 Kevin Dradjat , Massinissa Hamidi , Blaise Hanczar

Single-cell RNA sequencing (scRNA-seq) is essential for decoding tumor heterogeneity. However, pan-cancer research still faces two key challenges: learning discriminative and efficient single-cell representations, and establishing a…

基因组学 · 定量生物学 2025-12-04 Xiaoshui Huang , Tianlin Zhu , Yifan Zuo , Xue Xia , Zonghan Wu , Jiebin Yan , Dingli Hua , Zongyi Xu , Yuming Fang , Jian Zhang

Biological data mainly comprises of Deoxyribonucleic acid (DNA) and protein sequences. These are the biomolecules which are present in all cells of human beings. Due to the self-replicating property of DNA, it is a key constitute of genetic…

其他定量生物学 · 定量生物学 2020-06-04 Shakeela Bibi , Javed Iqbal , Adnan Iftekhar , Mir Hassan

Deep learning architectures such as convolutional neural networks and Transformers have revolutionized biological sequence modeling, with recent advances driven by scaling up foundation and task-specific models. The computational resources…

机器学习 · 计算机科学 2025-03-21 Krithik Ramesh , Sameed M. Siddiqui , Albert Gu , Michael D. Mitzenmacher , Pardis C. Sabeti

The application of deep learning methods, particularly foundation models, in biological research has surged in recent years. These models can be text-based or trained on underlying biological data, especially omics data of various types.…

人工智能 · 计算机科学 2024-12-06 Yoav Kan-Tor , Michael Morris Danziger , Eden Zohar , Matan Ninio , Yishai Shimoni

Tokenization sits at the boundary between high-throughput genomic input and GPU compute, posing challenges in both algorithm design and system throughput. Overlapping k-mer tokenization can introduce information leakage under masked…

基因组学 · 定量生物学 2026-01-12 Eliatan Niktab , Hardip Patel

The interactions between DNA, RNA, and proteins are fundamental to biological processes, as illustrated by the central dogma of molecular biology. Although modern biological pre-trained models have achieved great success in analyzing these…

机器学习 · 计算机科学 2025-12-02 Zicheng Liu , Siyuan Li , Zhiyuan Chen , Chang Yu , Qirong Yang , Yucheng Guo , Yujie Yang , Xiaoming Zhang , Stan Z. Li

Statistical inference on the cancer-site specificities of collective ultra-rare whole genome somatic mutations is an open problem. Traditional statistical methods cannot handle whole-genome mutation data due to their…

统计方法学 · 统计学 2023-01-02 Saptarshi Chakraborty , Zoe Guan , Colin B. Begg , Ronglai Shen

Functional annotation of microbial genomes is often biased toward protein-coding genes, leaving a vast, unexplored landscape of non-coding RNAs (ncRNAs) that are critical for regulating bacterial and archaeal physiology, stress response and…

基因组学 · 定量生物学 2025-07-17 Lauren Lui , Torben Nielsen

Recurrent Neural Networks (RNNs) with attention mechanisms have obtained state-of-the-art results for many sequence processing tasks. Most of these models use a simple form of encoder with attention that looks over the entire sequence and…

Sequence data, such as DNA, RNA, and protein sequences, exhibit intricate, multi-scale structures that pose significant challenges for conventional analysis methods, particularly those relying on alignment or purely statistical…

基因组学 · 定量生物学 2025-10-22 Jian Liu , Li Shen , Mushal Zia , Guo-Wei Wei

Advancements in genomic research such as high-throughput sequencing techniques have driven modern genomic studies into "big data" disciplines. This data explosion is constantly challenging conventional methods used in genomics. In parallel…

基因组学 · 定量生物学 2023-10-06 Tianwei Yue , Yuanxin Wang , Longxiang Zhang , Chunming Gu , Haoru Xue , Wenping Wang , Qi Lyu , Yujie Dun

Although theories regarding the role of sequence-specific DNA resonance in biology have abounded for over 40 years, the published evidence for it is lacking. Here, the authors reasoned that for sustained resonance signaling, the number of…

其他定量生物学 · 定量生物学 2019-10-28 Ivan Savelev , Max Myakishev-Rempel

High-quality training datasets are crucial for the development of effective protein design models, but existing synthetic datasets often include unfavorable sequence-structure pairs, impairing generative model performance. We leverage…