English
Related papers

Related papers: AbBiBench: A Benchmark for Antibody Binding Affini…

200 papers

While recent audio-visual models have demonstrated impressive performance, their robustness to distributional shifts at test-time remains not fully understood. Existing robustness benchmarks mainly focus on single modalities, making them…

While DeepMind has tentatively solved protein folding, its inverse problem -- protein design which predicts protein sequences from their 3D structures -- still faces significant challenges. Particularly, the lack of large-scale standardized…

Quantitative Methods · Quantitative Biology 2022-02-15 Zhangyang Gao , Cheng Tan , Stan Z. Li

As language models (LMs) become capable of handling a wide range of tasks, their evaluation is becoming as challenging as their development. Most generation benchmarks currently assess LMs using abstract evaluation criteria like helpfulness…

Despite the central role that antibodies play in the adaptive immune system and in biotechnology, much remains unknown about the quantitative relationship between an antibody's amino acid sequence and its antigen binding affinity. Here we…

Quantitative Methods · Quantitative Biology 2018-04-16 Rhys M. Adams , Thierry Mora , Aleksandra M. Walczak , Justin B. Kinney

Recent years have witnessed a surge in the development of protein foundation models, significantly improving performance in protein prediction and generative tasks ranging from 3D structure prediction and protein design to conformational…

Quantitative Methods · Quantitative Biology 2024-10-08 Fei Ye , Zaixiang Zheng , Dongyu Xue , Yuning Shen , Lihao Wang , Yiming Ma , Yan Wang , Xinyou Wang , Xiangxin Zhou , Quanquan Gu

We present TerraBind, a foundation model for protein-ligand structure and binding affinity prediction that achieves 26-fold faster inference than state-of-the-art methods while improving affinity prediction accuracy by $\sim$20\%. Current…

The accurate screening of candidate drug ligands against target proteins through computational approaches is of prime interest to drug development efforts. Such virtual screening depends in part on methods to predict the binding affinity…

Machine Learning · Computer Science 2024-10-22 Ho-Joon Lee , Prashant S. Emani , Mark B. Gerstein

The Adversarial Humanities Benchmark (AHB) evaluates whether model safety refusals survive a shift away from familiar harmful prompt forms. Starting from harmful tasks drawn from MLCommons AILuminate, the benchmark rewrites the same…

We introduce BikeBench, an engineering design benchmark for evaluating generative models on problems with multiple real-world objectives and constraints. As generative AI's reach continues to grow, evaluating its capability to understand…

Computational Engineering, Finance, and Science · Computer Science 2025-10-28 Lyle Regenwetter , Yazan Abu Obaideh , Fabien Chiotti , Ioanna Lykourentzou , Faez Ahmed

Data-driven generative models have emerged as promising approaches towards achieving efficient mechanical inverse design. However, due to prohibitively high cost in time and money, there is still lack of open-source and large-scale…

Computational Engineering, Finance, and Science · Computer Science 2024-10-29 Jian Liu , Jianyu Wu , Hairun Xie , Guoqing Zhang , Jing Wang , Wei Liu , Wanli Ouyang , Junjun Jiang , Xianming Liu , Shixiang Tang , Miao Zhang

Benchmarks are the de facto standard for tracking progress in large language models (LLMs), yet static test sets can rapidly saturate, become vulnerable to contamination, and are costly to refresh. Scalable evaluation of open-ended items…

Computation and Language · Computer Science 2026-03-24 Yandan Zheng , Haoran Luo , Zhenghong Lin , Wenjin Liu , Luu Anh Tuan

The evolution of Large Language Models (LLMs) into autonomous agents has expanded the scope of AI coding from localized code generation to complex, repository-level, and execution-driven problem solving. However, current benchmarks…

Software Engineering · Computer Science 2026-01-19 Jie Yang , Honglin Guo , Li Ji , Jiazheng Zhou , Rui Zheng , Zhikai Lei , Shuo Zhang , Zhiheng Xi , Shichun Liu , Yuxin Wang , Bo Wang , Yining Zheng , Tao Gui , Xipeng Qiu

We analyze the interactions between division, mutation and selection in a simplified evolutionary model, assuming that the population observed can be classified into fitness levels. The construction of our mathematical framework is…

Probability · Mathematics 2018-01-04 Irene Balelli , Vuk Milišić , Gilles Wainrib

Generative models can now propose thousands of \emph{de novo} antibody sequences, yet translating these designs into viable therapeutics remains constrained by the cost of biophysical characterization. Here we present CrossAbSense, a…

Biomolecules · Quantitative Biology 2026-04-13 Simon J. Crouzet

Automated ASD screening tools remain limited by single-architecture evaluations, axis-restricted assessment, and near-exclusive focus on adult cohorts, obscuring age-specific diagnostic patterns critical for early intervention. We introduce…

Machine Learning · Computer Science 2026-05-13 Shubhankit Singh , Hassan Shaikh , Kuldeep Raghuwanshi , Keshav Bulia

The adaptive immune response, largely mediated by B-cell receptors (BCRs), plays a crucial role for effective pathogen neutralization due to its diversity and antigen specificity. Designing BCRs de novo, or from scratch, has been…

Biomolecules · Quantitative Biology 2024-09-11 Desmond Kuan , Amir Barati Farimani

Nucleotide sequence variation can induce significant shifts in functional fitness. Recent nucleotide foundation models promise to predict such fitness effects directly from sequence, yet heterogeneous datasets and inconsistent preprocessing…

Genomics · Quantitative Biology 2025-11-06 Zhongmin Li , Runze Ma , Jiahao Tan , Chengzi Tan , Shuangjia Zheng

Antibodies are proteins produced by the immune system that can identify and neutralise a wide variety of antigens with high specificity and affinity, and constitute the most successful class of biotherapeutics. With the advent of…

Biomolecules · Quantitative Biology 2024-03-27 Henry Kenlay , Frédéric A. Dreyer , Aleksandr Kovaltsuk , Dom Miketa , Douglas Pires , Charlotte M. Deane

We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations. Existing health benchmarks emphasize medical question answering, broad health…

Currently, the field of structure-based drug design is dominated by three main types of algorithms: search-based algorithms, deep generative models, and reinforcement learning. While existing works have typically focused on comparing models…

Machine Learning · Computer Science 2026-01-22 Kangyu Zheng , Kai Zhang , Jiale Tan , Xuehan Chen , Yingzhou Lu , Zaixi Zhang , Lichao Sun , Marinka Zitnik , Tianfan Fu , Zhiding Liang