English
Related papers

Related papers: Mass spectrometry based protein identification wit…

200 papers

Recent advances in topology-based modeling have accelerated progress in physical modeling and molecular studies, including applications to protein-ligand binding affinity. In this work, we introduce the Persistent Laplacian Decision Tree…

Biomolecules · Quantitative Biology 2024-12-25 Xingjian Xu , Jiahui Chen , Chunmei Wang

In recent years, there has been an explosion of research on the application of deep learning to the prediction of various peptide properties, due to the significant development and market potential of peptides. Molecular dynamics has…

Biomolecules · Quantitative Biology 2023-07-19 Zihan Liu , Jiaqi Wang , Yun Luo , Shuang Zhao , Wenbin Li , Stan Z. Li

Leveraging multiple training datasets to scale up image segmentation models is beneficial for increasing robustness and semantic understanding. Individual datasets have well-defined ground truth with non-overlapping mask layouts and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Qilong Zhangli , Di Liu , Abhishek Aich , Dimitris Metaxas , Samuel Schulter

Proteins are biomolecules of life. They fold into a great variety of three-dimensional (3D) shapes. Underlying these folding patterns are many recurrent structural fragments or building blocks (analogous to `LEGO bricks'). This paper…

Quantitative Methods · Quantitative Biology 2013-10-08 Arun S. Konagurthu , Arthur M. Lesk , David Abramson , Peter J. Stuckey , Lloyd Allison

Deep neural-network-based language models (LMs) are increasingly applied to large-scale protein sequence data to predict protein function. However, being largely black-box models and thus challenging to interpret, current protein LM…

Quantitative Methods · Quantitative Biology 2024-08-06 Mai Ha Vu , Rahmad Akbar , Philippe A. Robert , Bartlomiej Swiatczak , Victor Greiff , Geir Kjetil Sandve , Dag Trygve Truslew Haug

Numerous machine learning (ML) models employed in protein function and structure prediction depend on evolutionary information, which is captured through multiple-sequence alignments (MSA) or position-specific scoring matrices (PSSM) as…

Computational Engineering, Finance, and Science · Computer Science 2024-01-17 Issar Arab

The identification of essential proteins can help in understanding the minimum requirements for cell survival and development. Network-based centrality approaches are commonly used to identify essential proteins from protein-protein…

Computational Engineering, Finance, and Science · Computer Science 2023-11-13 Li Pan , Haoyue Wang , Jing Sun , Bin Li , Bo Yang , Wenbin Li

Proteomics is the large scale study of protein structure and function from biological systems through protein identification and quantification. "Shotgun proteomics" or "bottom-up proteomics" is the prevailing strategy, in which proteins…

Motivation Protein fold recognition is an important problem in structural bioinformatics. Almost all traditional fold recognition methods use sequence (homology) comparison to indirectly predict the fold of a tar get protein based on the…

Machine Learning · Computer Science 2017-06-06 Jie Hou , Badri Adhikari , Jianlin Cheng

Computational protein design is experiencing a transformation driven by AI/ML. However, the range of potential protein sequences and structures is astronomically vast, even for moderately sized proteins. Hence, achieving convergence between…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-10-09 Aymen Alsaadi , Jonathan Ash , Mikhail Titov , Matteo Turilli , Andre Merzky , Shantenu Jha , Sagar Khare

High throughput screening of compounds (chemicals) is an essential part of drug discovery [7], involving thousands to millions of compounds, with the purpose of identifying candidate hits. Most statistical tools, including the industry…

Machine Learning · Statistics 2017-09-29 Ivo D. Shterev , David B. Dunson , Cliburn Chan , Gregory D. Sempowski

Missing data present challenges in data analysis. Naive analyses such as complete-case and available-case analysis may introduce bias and loss of efficiency, and produce unreliable results. Multiple imputation (MI) is one of the most widely…

Methodology · Statistics 2019-05-15 Domonique W. Hodge , Sandra E. Safo , Qi Long

In experimental nuclear and particle physics, the extraction of high-purity samples of rare events critically depends on the efficiency and accuracy of particle identification (PID). In this work, we present a PID method applied to HADES…

Data Analysis, Statistics and Probability · Physics 2025-11-18 Marvin Kohls

While many good textbooks are available on Protein Structure, Molecular Simulations, Thermodynamics and Bioinformatics methods in general, there is no good introductory level book for the field of Structural Bioinformatics. This book aims…

Determination of binding affinity of proteins in the formation of protein complexes requires sophisticated, expensive and time-consuming experimentation which can be replaced with computational methods. Most computational prediction…

Quantitative Methods · Quantitative Biology 2020-12-14 Wajid Arshad Abbasi , Fahad Ul Hassan , Adiba Yaseen , Fayyaz Ul Amir Afsar Minhas

Motivation: Drug discovery demands rapid quantification of compound-protein interaction (CPI). However, there is a lack of methods that can predict compound-protein affinity from sequences alone with high applicability, accuracy, and…

Biomolecules · Quantitative Biology 2020-12-17 Mostafa Karimi , Di Wu , Zhangyang Wang , Yang Shen

Protein sequence classification involves feature selection for accurate classification. Popular protein sequence classification techniques involve extraction of specific features from the sequences. Researchers apply some well-known…

Computational Engineering, Finance, and Science · Computer Science 2012-11-21 Suprativ Saha , Rituparna Chaki

Background: Coevolution within a protein family is often predicted using statistics that measure the degree of covariation between positions in the protein sequence. Mutual Information is a measure of dependence between two random variables…

Populations and Evolution · Quantitative Biology 2013-04-17 Russell J. Dickson , Gregory B. Gloor

Background: Missing data is a common challenge in mass spectrometry-based metabolomics, which can lead to biased and incomplete analyses. The integration of whole-genome sequencing (WGS) data with metabolomics data has emerged as a…

During the last decade, a large number of different numerical methods have been proposed to tackle the automatic identification and quantification in {\gamma}-ray spectrometry. However, the lack of common benchmarks, including datasets,…

Machine Learning · Computer Science 2025-08-13 Dinh Triem Phan , Jérôme Bobin , Cheick Thiam , Christophe Bobin