English
Related papers

Related papers: An ISA-Tab specification for protein titration dat…

200 papers

Topological data analysis (TDA) detects geometric structure in biological data. However, many TDA algorithms are memory intensive and impractical for massive datasets. Here, we introduce a statistical protocol that reduces TDA's memory…

Quantitative Methods · Quantitative Biology 2025-09-05 Andrew J. Stier , Naichen Shi , Raed Al Kontar , Chad Giusti , Marc G. Berman

Bioactivity data plays a key role in drug discovery and repurposing. The resource-demanding nature of \textit{in vitro} and \textit{in vivo} experiments, as well as the recent advances in data-driven computational biochemistry research,…

We introduce a new discriminant analysis method (Empirical Discriminant Analysis or EDA) for binary classification in machine learning. Given a dataset of feature vectors, this method defines an empirical feature map transforming the…

Machine Learning · Statistics 2012-10-30 Mark A. Kon , Nikolay Nikolaev

Statistical agencies frequently release frequency tables derived from microdata, but small frequency cells may lead to disclosure risks. We present \texttt{iLBA}, an open-source \textsf{R} package for confidential dissemination of…

Computation · Statistics 2026-04-07 Jeehyun Hwang , Dongsun Yoon , Sungkyu Jung , Min-Jeong Park , Inkwon Yeo

Motivation: Human cancer is caused by the accumulation of somatic mutations in tumor suppressors and oncogenes within the genome. In the case of oncogenes, recent theory suggests that there are only a few key "driver" mutations responsible…

Genomics · Quantitative Biology 2013-07-16 Gregory Ryslik , Yuwei Cheng , Kei-Hoi Cheung , Yorgo Modis , Hongyu Zhao

This report investigates Training Data Attribution (TDA) and its potential importance to and tractability for reducing extreme risks from AI. First, we discuss the plausibility and amount of effort it would take to bring existing TDA…

Computers and Society · Computer Science 2025-01-23 Deric Cheng , Juhan Bae , Justin Bullock , David Kristofferson

We introduce a new, simplified model of proteins, which we call protein metastructure. The metastructure of a protein carries information about its secondary structure and $\beta$-strand conformations. Furthermore, protein metastructure…

Biomolecules · Quantitative Biology 2021-11-30 Jørgen Ellegaard Andersen , Hiroyuki Fuji , Yuki Koyanagi

This review discusses research developments and applications of isotachophoresis (ITP) to the initiation, control, and acceleration of chemical reactions, emphasizing reactions involving biomolecular reactants such as nucleic acids,…

Biomolecules · Quantitative Biology 2017-08-29 Charbel Eid , Juan G. Santiago

Many questions in computational social science rely on datasets assembled from heterogeneous online sources, a process that is often labor-intensive, costly, and difficult to reproduce. Recent advances in large language models enable…

Computation and Language · Computer Science 2026-01-07 Mengyi Sun

This paper describes the CAVES attestation protocol and presents a tool-supported analysis showing that the runs of the protocol achieve stated goals. The goals are stated formally by annotating the protocol with logical formulas using the…

Cryptography and Security · Computer Science 2012-07-03 John D. Ramsdell , Joshua D. Guttman , Jonathan K. Millen , Brian O'Hanlon

Multiple sequence alignment (MSA) data play a crucial role in the study of protein mutations, with contact prediction being a notable application. Existing methods are often model-based or algorithmic and typically do not incorporate…

Methodology · Statistics 2026-01-23 Fan Yang , Zhao Ren , Wen Zhou , Kejue Jia , Robert Jernigan

Protein-protein interaction networks provide a graph-level view of cellular organization, yet their functional modules are overlapping, noisy, and difficult to interpret from cluster assignments alone. Existing community-detection methods…

Social and Information Networks · Computer Science 2026-05-21 Sima Soltani , Mehrdad Jalali , Yahya Forghani

Identifying similar protein sequences is a core step in many computational biology pipelines such as detection of homologous protein sequences, generation of similarity protein graphs for downstream analysis, functional annotation and gene…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-10-01 Oguz Selvitopi , Saliya Ekanayake , Giulia Guidi , Georgios Pavlopoulos , Ariful Azad , Aydin Buluc

Protein quality assessment (QA) by ranking and selecting protein models has long been viewed as one of the major challenges for protein tertiary structure prediction. Especially, estimating the quality of a single protein model, which is…

Artificial Intelligence · Computer Science 2016-07-18 Renzhi Cao , Debswapna Bhattacharya , Jie Hou , Jianlin Cheng

We participated in three of the protein-protein interaction subtasks of the Second BioCreative Challenge: classification of abstracts relevant for protein-protein interaction (IAS), discovery of protein pairs (IPS) and text passages…

Calibration refers to the estimation of unknown parameters which are present in computer experiments but not available in physical experiments. An accurate estimation of these parameters is important because it provides a scientific…

Methodology · Statistics 2019-03-21 Chih-Li Sung , Ying Hung , William Rittase , Cheng Zhu , C. F. Jeff Wu

Tabular data plays a pivotal role in various fields, making it a popular format for data manipulation and exchange, particularly on the web. The interpretation, extraction, and processing of tabular information are invaluable for…

Artificial Intelligence · Computer Science 2024-11-20 Marco Cremaschi , Blerina Spahiu , Matteo Palmonari , Ernesto Jimenez-Ruiz

The primary structure of proteins, that is their sequence, represents one of the most abundant set of experimental data concerning biomolecules. The study of correlations in families of co--evolving proteins by means of an inverse…

Biomolecules · Quantitative Biology 2015-06-16 Sara Lui , Guido Tiana

The information retrieval (IR) community has a strong tradition of making the computational artifacts and resources available for future reuse, allowing the validation of experimental results. Besides the actual test collections, the…

Information Retrieval · Computer Science 2022-07-20 Timo Breuer , Jüri Keller , Philipp Schaer

Knowledge distillation uses both real hard labels and soft labels predicted by teacher models as supervision. Intuitively, we expect the soft labels and hard labels to be concordant w.r.t. their orders of probabilities. However, we found…

Machine Learning · Computer Science 2021-07-07 Wanyun Cui , Sen Yan
‹ Prev 1 4 5 6 7 8 10 Next ›