English
Related papers

Related papers: MLReplicate: Benchmarking Autonomous Research Syst…

200 papers

Benchmarks are the de facto standard for tracking progress in large language models (LLMs), yet static test sets can rapidly saturate, become vulnerable to contamination, and are costly to refresh. Scalable evaluation of open-ended items…

Computation and Language · Computer Science 2026-03-24 Yandan Zheng , Haoran Luo , Zhenghong Lin , Wenjin Liu , Luu Anh Tuan

With the recent advances in A.I. methodologies and their application to medical imaging, there has been an explosion of related research programs utilizing these techniques to produce state-of-the-art classification performance. Ultimately,…

As scientific knowledge grows at an unprecedented pace, evaluation benchmarks must evolve to reflect new discoveries and ensure language models are tested on current, diverse literature. We propose a scalable, modular framework for…

Computation and Language · Computer Science 2025-09-16 Ozan Gokdemir , Neil Getty , Robert Underwood , Sandeep Madireddy , Franck Cappello , Arvind Ramanathan , Ian T. Foster , Rick L. Stevens

How many mistakes do published AI papers contain? Peer-reviewed publications form the foundation upon which new research and knowledge are built. Errors that persist in the literature can propagate unnoticed, creating confusion in follow-up…

Artificial Intelligence · Computer Science 2025-12-08 Federico Bianchi , Yongchan Kwon , Zachary Izzo , Linjun Zhang , James Zou

Given that Large Language Models (LLMs) have made significant progress in writing code, can they now be used to autonomously reproduce results from research repositories? Such a capability would be a boon to the research community, helping…

Artificial Intelligence · Computer Science 2024-09-12 Ben Bogin , Kejuan Yang , Shashank Gupta , Kyle Richardson , Erin Bransom , Peter Clark , Ashish Sabharwal , Tushar Khot

We present a novel platform for evaluating the capability of Large Language Models (LLMs) to autonomously compose and critique survey papers spanning a vast array of disciplines including sciences, humanities, education, and law. Within…

Computation and Language · Computer Science 2023-10-11 Thanh Gia Hieu Khuong , Benedictus Kent Rachmat

Increasingly, artificial intelligence (AI) and machine learning (ML) are used in eScience applications [9]. While these approaches have great potential, the literature has shown that ML-based approaches frequently suffer from results that…

Machine Learning · Computer Science 2024-07-03 Zhiwei Li , Carl Kesselman , Mike D'Arch , Michael Pazzani , Benjamin Yizing Xu

Recently, multimodal large language models (MLLMs) have achieved significant advancements across various domains, and corresponding evaluation benchmarks have been continuously refined and improved. In this process, benchmarks in the…

Computation and Language · Computer Science 2025-08-20 Jiacheng Ruan , Dan Jiang , Xian Gao , Ting Liu , Yuzhuo Fu , Yangyang Kang

Machine learning (ML) has become a vital part in many aspects of our daily life. However, building well performing machine learning applications requires highly specialized data scientists and domain experts. Automated machine learning…

Machine Learning · Computer Science 2021-01-27 Marc-André Zöller , Marco F. Huber

Machine learning algorithms designed to characterize, monitor, and intervene on human health (ML4H) are expected to perform safely and reliably when operating at scale, potentially outside strict human supervision. This requirement warrants…

Machine Learning · Computer Science 2019-07-03 Matthew B. A. McDermott , Shirly Wang , Nikki Marinsek , Rajesh Ranganath , Marzyeh Ghassemi , Luca Foschini

With the rapid progress of multimodal large language models (MLLMs), AI already performs well at literature retrieval and certain reasoning tasks, serving as a capable assistant to human researchers, yet it remains far from autonomous…

Artificial Intelligence · Computer Science 2026-03-31 Rongjin Li , Zichen Tang , Xianghe Wang , Xinyi Hu , Zhengyu Wang , Zhengyu Lu , Yiling Huang , Jiayuan Chen , Weisheng Tan , Jiacheng Liu , Zhongjun Yang , Haihong E

Learned representations of scientific documents can serve as valuable input features for downstream tasks without further fine-tuning. However, existing benchmarks for evaluating these representations fail to capture the diversity of…

Computation and Language · Computer Science 2023-11-14 Amanpreet Singh , Mike D'Arcy , Arman Cohan , Doug Downey , Sergey Feldman

As AI systems enter high-stakes domains, evaluation must extend beyond predictive accuracy to include explainability, fairness, robustness, and sustainability. We introduce RAISE (Responsible AI Scoring and Evaluation), a unified framework…

Machine Learning · Computer Science 2025-10-22 Loc Phuc Truong Nguyen , Hung Thanh Do

Information retrieval (IR) evaluation remains challenging due to incomplete IR benchmark datasets that contain unlabeled relevant chunks. While LLMs and LLM-human hybrid strategies reduce costly human effort, they remain prone to LLM…

Computation and Language · Computer Science 2026-02-09 Minjeong Ban , Jeonghwan Choi , Hyangsuk Min , Nicole Hee-Yeon Kim , Minseok Kim , Jae-Gil Lee , Hwanjun Song

Two goals - improving replicability and accountability of Machine Learning research respectively, have accrued much attention from the AI ethics and the Machine Learning community. Despite sharing the measures of improving transparency, the…

Computers and Society · Computer Science 2025-08-14 Tianqi Kou

Reproducibility remains a central challenge in machine learning (ML), especially in collaborative eScience projects where teams iterate over data, features, and models. Current ML workflows are often dynamic yet fragmented, relying on…

Machine Learning · Computer Science 2025-06-23 Zhiwei Li , Carl Kesselman , Tran Huy Nguyen , Benjamin Yixing Xu , Kyle Bolo , Kimberley Yu

Peer review, the bedrock of scientific advancement in machine learning (ML), is strained by a crisis of scale. Exponential growth in manuscript submissions to premier ML venues such as NeurIPS, ICML, and ICLR is outpacing the finite…

Artificial Intelligence · Computer Science 2025-06-30 Qiyao Wei , Samuel Holt , Jing Yang , Markus Wulfmeier , Mihaela van der Schaar

Peer review is a critical component of scientific progress in the fields like AI, but the rapid increase in submission volume has strained the reviewing system, which inevitably leads to reviewer shortages and declines review quality.…

Computation and Language · Computer Science 2026-03-16 Daoze Zhang , Zhijian Bao , Sihang Du , Zhiyi Zhao , Kuangling Zhang , Dezheng Bao , Yang Yang

As qualitative researchers show growing interest in using automated tools to support interpretive analysis, a large language model (LLM) is often introduced into an analytic workflow as is, without systematic evaluation of interpretive…

Computation and Language · Computer Science 2026-04-02 Songhee Han , Jueun Shin , Jiyoon Han , Bung-Woo Jun , Hilal Ayan Karabatman

The rapid development of science and technology has been accompanied by an exponential growth in peer-reviewed scientific publications. At the same time, the review of each paper is a laborious process that must be carried out by subject…

Computation and Language · Computer Science 2021-02-02 Weizhe Yuan , Pengfei Liu , Graham Neubig