中文
相关论文

相关论文: Does AI for science need another ImageNet Or total…

200 篇论文

Forecasting is a challenging task that offers a clearly measurable way to study AI systems. Forecasting requires a large amount of research on the internet, and evaluations require time for events to happen, making the development of…

Machine-learned force fields (MLFFs), especially pre-trained foundation models, are transforming computational materials science by enabling ab initio-level accuracy at molecular dynamics scales. Yet their rapid rise raises a key question:…

化学物理 · 物理学 2025-10-20 Yi Cao , Paulette Clancy

Autonomous scientific research is significantly advanced thanks to the development of AI agents. One key step in this process is finding the right scientific literature, whether to explore existing knowledge for a research problem, or to…

Markets are a promising way to coordinate AI agent activity for similar reasons to those used to justify markets more broadly. In order to effectively participate in markets, agents need to have informative signals of their own ability to…

人工智能 · 计算机科学 2026-04-28 Andrey Fradkin , Rohit Krishnan

A force field as accurate as quantum mechanics (QM) and as fast as molecular mechanics (MM), with which one can simulate a biomolecular system efficiently enough and meaningfully enough to get quantitative insights, is among the most ardent…

There is widespread optimism that frontier Large Language Models (LLMs) and LLM-augmented systems have the potential to rapidly accelerate scientific discovery across disciplines. Today, many benchmarks exist to measure LLM knowledge and…

We introduce Meta MLGym and MLGym-Bench, a new framework and benchmark for evaluating and developing LLM agents on AI research tasks. This is the first Gym environment for machine learning (ML) tasks, enabling research on reinforcement…

AI research agents are demonstrating great potential to accelerate scientific progress by automating the design, implementation, and training of machine learning models. We focus on methods for improving agents' performance on MLE-bench, a…

As large language models (LLMs) transition from general knowledge retrieval to complex scientific discovery, their evaluation standards must also incorporate the rigorous norms of scientific inquiry. Existing benchmarks exhibit a critical…

With advances in digital technology, the classification of medical images has become a crucial step for image-based clinical decision support systems. Automatic medical image classification represents a pivotal domain where the use of AI…

计算机视觉与模式识别 · 计算机科学 2024-09-09 Abu Adnan Sadi , Labib Chowdhury , Nusrat Jahan , Mohammad Newaz Sharif Rafi , Radeya Chowdhury , Faisal Ahamed Khan , Nabeel Mohammed

Benchmarks are essential for unified evaluation and reproducibility. The rapid rise of Artificial Intelligence for Software Engineering (AI4SE) has produced numerous benchmarks for tasks such as code generation and bug repair. However, this…

软件工程 · 计算机科学 2025-12-15 Roham Koohestani , Philippe de Bekker , Begüm Koç , Maliheh Izadi

Extreme-edge scientific applications use machine learning models to analyze sensor data and make real-time decisions. Their stringent latency and throughput requirements demand small batch sizes and require that model weights remain fully…

硬件体系结构 · 计算机科学 2026-04-22 Zhenghua Ma , G Abarajithan , Dimitrios Danopoulos , Olivia Weng , Francesco Restuccia , Ryan Kastner

Current benchmarks that test LLMs on static, already-solved problems (e.g., math word problems) effectively demonstrated basic capability acquisition. The natural progression has been toward larger, more comprehensive and challenging…

机器学习 · 计算机科学 2025-12-15 Alwin Jin , Sean M. Hendryx , Vaskar Nath

AI agents hold the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new directions of inquiry; indeed, there are now many such agents, ranging…

LLMs demand significant computational resources for both pre-training and fine-tuning, requiring distributed computing capabilities due to their large model sizes \cite{sastry2024computing}. Their complex architecture poses challenges…

分布式、并行与集群计算 · 计算机科学 2024-12-03 Todor Ivanov , Valeri Penchev

Machine learning force fields (MLFF) have been proposed to accelerate molecular dynamics (MD) simulation, which finds widespread applications in chemistry and biomedical research. Even for the most data-efficient MLFFs, reaching chemical…

定量方法 · 定量生物学 2023-06-07 Alexander Bukharin , Tianyi Liu , Shengjie Wang , Simiao Zuo , Weihao Gao , Wen Yan , Tuo Zhao

Machine learning force fields (MLFFs) are a promising approach to balance the accuracy of quantum mechanics with the efficiency of classical potentials, yet selecting an optimal model amid increasingly diverse architectures that delivers…

机器学习 · 计算机科学 2025-12-09 Bangchen Yin , Yue Yin , Yuda W. Tang , Hai Xiao

Medical AI faces challenges in privacy-preserving collaborative learning while ensuring fairness across heterogeneous healthcare institutions. Current federated learning approaches suffer from static architectures, slow convergence (45-73…

计算机与社会 · 计算机科学 2025-10-24 Jahidul Arafat , Fariha Tasmin , Sanjaya Poudel , Iftekhar Haider

In this community review report, we discuss applications and techniques for fast machine learning (ML) in science -- the concept of integrating power ML methods into the real-time experimental data processing loop to accelerate scientific…

机器学习 · 计算机科学 2023-02-07 Allison McCarn Deiana , Nhan Tran , Joshua Agar , Michaela Blott , Giuseppe Di Guglielmo , Javier Duarte , Philip Harris , Scott Hauck , Mia Liu , Mark S. Neubauer , Jennifer Ngadiuba , Seda Ogrenci-Memik , Maurizio Pierini , Thea Aarrestad , Steffen Bahr , Jurgen Becker , Anne-Sophie Berthold , Richard J. Bonventre , Tomas E. Muller Bravo , Markus Diefenthaler , Zhen Dong , Nick Fritzsche , Amir Gholami , Ekaterina Govorkova , Kyle J Hazelwood , Christian Herwig , Babar Khan , Sehoon Kim , Thomas Klijnsma , Yaling Liu , Kin Ho Lo , Tri Nguyen , Gianantonio Pezzullo , Seyedramin Rasoulinezhad , Ryan A. Rivera , Kate Scholberg , Justin Selig , Sougata Sen , Dmitri Strukov , William Tang , Savannah Thais , Kai Lukas Unger , Ricardo Vilalta , Belinavon Krosigk , Thomas K. Warburton , Maria Acosta Flechas , Anthony Aportela , Thomas Calvet , Leonardo Cristella , Daniel Diaz , Caterina Doglioni , Maria Domenica Galati , Elham E Khoda , Farah Fahim , Davide Giri , Benjamin Hawks , Duc Hoang , Burt Holzman , Shih-Chieh Hsu , Sergo Jindariani , Iris Johnson , Raghav Kansal , Ryan Kastner , Erik Katsavounidis , Jeffrey Krupa , Pan Li , Sandeep Madireddy , Ethan Marx , Patrick McCormack , Andres Meza , Jovan Mitrevski , Mohammed Attia Mohammed , Farouk Mokhtar , Eric Moreno , Srishti Nagu , Rohin Narayan , Noah Palladino , Zhiqiang Que , Sang Eon Park , Subramanian Ramamoorthy , Dylan Rankin , Simon Rothman , Ashish Sharma , Sioni Summers , Pietro Vischia , Jean-Roch Vlimant , Olivia Weng