English
Related papers

Related papers: ProtiGeno: a prokaryotic short gene finder using p…

200 papers

Computational identification of promoters is notoriously difficult as human genes often have unique promoter sequences that provide regulation of transcription and interaction with transcription initiation complex. While there are many…

Genomics · Quantitative Biology 2018-10-04 Ramzan Umarov , Hiroyuki Kuwahara , Yu Li , Xin Gao , Victor Solovyev

Protein engineering seeks to identify protein sequences with optimized properties. When guided by machine learning, protein sequence generation methods can draw on prior knowledge and experimental efforts to improve this process. In this…

Quantitative Methods · Quantitative Biology 2021-05-28 Zachary Wu , Kadina E. Johnston , Frances H. Arnold , Kevin K. Yang

Large Language models (LLMs) have emerged as powerful tools for addressing challenges across diverse domains. Notably, recent studies have demonstrated that large language models significantly enhance the efficiency of biomolecular analysis…

Computation and Language · Computer Science 2025-03-07 Jiyue Jiang , Zikang Wang , Yuheng Shan , Heyan Chai , Jiayi Li , Zixian Ma , Xinrui Zhang , Yu Li

Tabular biomedical data poses challenges in machine learning because it is often high-dimensional and typically low-sample-size (HDLSS). Previous research has attempted to address these challenges via local feature selection, but existing…

Machine Learning · Computer Science 2024-06-04 Xiangjian Jiang , Andrei Margeloiu , Nikola Simidjievski , Mateja Jamnik

Accurately predicting protein fitness with minimal experimental data is a persistent challenge in protein engineering. We introduce PRIMO (PRotein In-context Mutation Oracle), a transformer-based framework that leverages in-context learning…

Biomolecules · Quantitative Biology 2025-12-03 Felix Teufel , Aaron W. Kollasch , Yining Huang , Ole Winther , Kevin K. Yang , Pascal Notin , Debora S. Marks

The high-throughput data generated by microarray experiments provides complete set of genes being expressed in a given cell or in an organism under particular conditions. The analysis of these enormous data has opened a new dimension for…

Computational Engineering, Finance, and Science · Computer Science 2012-11-12 Khalid Raza , Akhilesh Mishra

Proteolysis targeting chimeras (PROTACs) are small molecules that trigger the breakdown of traditionally ``undruggable'' proteins by binding simultaneously to their targets and degradation-associated proteins. A key challenge in their…

Biomolecules · Quantitative Biology 2024-05-14 Bo Qiang , Wenxian Shi , Yuxuan Song , Menghua Wu

Designing protein sequences that fold into a target 3D structure, known as protein inverse folding, is a fundamental challenge in protein engineering. While recent deep learning methods have achieved impressive performance by recovering…

Biomolecules · Quantitative Biology 2025-06-03 Mengdi Liu , Xiaoxue Cheng , Zhangyang Gao , Hong Chang , Cheng Tan , Shiguang Shan , Xilin Chen

Motivation: Microarray data has been recently been shown to be efficacious in distinguishing closely related cell types that often appear in the diagnosis of cancer. It is useful to determine the minimum number of genes needed to do such a…

Biological Physics · Physics 2007-05-23 J. M. Deutsch

Understanding biological processes, drug development, and biotechnological advancements requires a detailed analysis of protein structures and functions, a task that is inherently complex and time-consuming in traditional protein research.…

Artificial Intelligence · Computer Science 2025-04-21 Yijia Xiao , Edward Sun , Yiqiao Jin , Qifan Wang , Wei Wang

DNA sequence encoding is fundamental to gene function prediction, protein synthesis, and diverse downstream biological tasks. Despite the substantial progress achieved by large-scale DNA sequence pretraining, existing studies have…

Machine Learning · Computer Science 2026-04-21 Zhijiang Tang , Jiaxin Qi , Yan Cui , Jinli Ou , Yuhua Zheng , Jianqiang Huang

Data on the number of Open Reading Frames (ORFs) coded by genomes from the 3 domains of Life show some notable general features including essential differences between the Prokaryotes and Eukaryotes, with the number of ORFs growing linearly…

Genomics · Quantitative Biology 2012-05-31 James L. Friar , Terrance Goldman , Juan Pérez-Mercader

The characterization of drug-protein interactions is crucial in the high-throughput screening for drug discovery. The deep learning-based approaches have attracted attention because they can predict drug-protein interactions without…

Machine Learning · Computer Science 2020-12-22 QHwan Kim , Joon-Hyuk Ko , Sunghoon Kim , Nojun Park , Wonho Jhe

Next-generation sequencing technologies generate millions of short sequence reads, which are usually aligned to a reference genome. In many applications, the key information required for downstream analysis is the number of reads mapping to…

Genomics · Quantitative Biology 2016-07-26 Yang Liao , Gordon K Smyth , Wei Shi

Weight space learning aims to extract information about a neural network, such as its training dataset or generalization error. Recent approaches learn directly from model weights, but this presents many challenges as weights are…

Machine Learning · Computer Science 2025-10-23 Jonathan Kahana , Eliahu Horwitz , Imri Shuval , Yedid Hoshen

Studying the function of proteins is important for understanding the molecular mechanisms of life. The number of publicly available protein structures has increasingly become extremely large. Still, the determination of the function of a…

Machine Learning · Computer Science 2018-03-02 Wajdi Dhifli , Abdoulaye Baniré Diallo

Complete genome sequences contain valuable information about natural selection, but extracting this information for short, widely scattered noncoding elements remains a challenging problem. Here we introduce a new computational method for…

Genomics · Quantitative Biology 2015-03-19 Ilan Gronau , Leonardo Arbiza , Jaaved Mohammed , Adam Siepel

The subcellular location of a protein can provide valuable information about its function. With the rapid increase of sequenced genomic data, the need for an automated and accurate tool to predict subcellular localization becomes…

Neural and Evolutionary Computing · Computer Science 2007-10-12 Sabu M. Thampi , K. Chandra Sekaran

Self-supervised protein language models have proved their effectiveness in learning the proteins representations. With the increasing computational power, current protein language models pre-trained with millions of diverse sequences can…

Biomolecules · Quantitative Biology 2022-11-02 Ningyu Zhang , Zhen Bi , Xiaozhuan Liang , Siyuan Cheng , Haosen Hong , Shumin Deng , Jiazhang Lian , Qiang Zhang , Huajun Chen

Short-read DNA sequencing instruments can yield over 1e+12 bases per run, typically composed of reads 150 bases long. Despite this high throughput, de novo assembly algorithms have difficulty reconstructing contiguous genome sequences using…

Genomics · Quantitative Biology 2023-06-09 Eric Chen , Justin Chu , Jessica Zhang , Rene L. Warren , Inanc Birol