English
Related papers

Related papers: gwBenchmarks: Stress-Testing LLM Agents on High-Pr…

200 papers

The phenomenon of Gravitational Wave (GW) analysis has grown in popularity as technology has advanced and the process of observing gravitational waves has become more precise. Although the sensitivity and the frequency of observation of GW…

Machine Learning · Computer Science 2023-11-07 Elena-Simona Apostol , Ciprian-Octavian Truică

Standard benchmarks fixate on how well large language model (LLM) agents perform in finance, yet say little about whether they are safe to deploy. We argue that accuracy metrics and return-based scores provide an illusion of reliability,…

General Finance · Quantitative Finance 2025-06-03 Zichen Chen , Jiaao Chen , Jianda Chen , Misha Sra

Physics simulators are essential in science and engineering, enabling the analysis, control, and design of complex systems. In experimental sciences, they are increasingly used to automate experimental design, often via combinatorial search…

Large Language Models (LLMs), such as GPT-4, have demonstrated impressive mathematical reasoning capabilities, achieving near-perfect performance on benchmarks like GSM8K. However, their application in personalized education remains limited…

Computer Vision and Pattern Recognition · Computer Science 2025-02-20 Yi-Fan Zhang , Hang Li , Dingjie Song , Lichao Sun , Tianlong Xu , Qingsong Wen

Binary black hole (BBH) mergers detected via gravitational waves are addressing key open questions in astrophysics, cosmology, and fundamental physics. Our scientific conclusions rely on extracting accurate source parameters, for which we…

General Relativity and Quantum Cosmology · Physics 2026-03-30 Parthapratim Mahapatra , Jonathan E. Thompson , Edward Fauchon-Jones , Mark Hannam

Multimodal Large Language Models (MLLMs) are evolving from passive observers into active agents, solving problems through Visual Expansion (invoking visual tools) and Knowledge Expansion (open-web search). However, existing evaluations fall…

Artificial Intelligence · Computer Science 2026-04-06 Qianshan Wei , Yishan Yang , Siyi Wang , Jinglin Chen , Binyu Wang , Jiaming Wang , Shuang Chen , Zechen Li , Yang Shi , Yuqi Tang , Weining Wang , Yi Yu , Chaoyou Fu , Qi Li , Yi-Fan Zhang

The advances made by Large Language Models (LLMs) have led to the pursuit of LLM agents that can solve intricate, multi-step reasoning tasks. As with any research pursuit, benchmarking and evaluation are key corner stones to efficient and…

Artificial Intelligence · Computer Science 2024-04-10 Luca Gioacchini , Giuseppe Siracusano , Davide Sanvito , Kiril Gashteovski , David Friede , Roberto Bifulco , Carolin Lawrence

Astrometric measurements provide a unique avenue for constraining the stochastic gravitational wave background (SGWB). In this work, we investigate the application of two neural network architectures, a fully connected network and a graph…

Cosmology and Nongalactic Astrophysics · Physics 2026-02-18 Marienza Caldarola , Gonzalo Morrás , Santiago Jaraba , Sachiko Kuroyanagi , Savvas Nesseris , Juan García-Bellido

Vision Language Models (VLMs) have been applied to several specific domains and have shown strong problem-solving capabilities. However, astronomical imaging, a quite complex problem involving multidisciplinary knowledge and several…

Multiagent Systems · Computer Science 2026-04-20 Yaohui Han , Tianshuo Wang , Zixi Zhao , Zhengchun Zhu , Shuo Ren , Yiru Wang , Rongliang Fu , Tinghuan Chen , Tsung-Yi Ho

Recent work in scientific machine learning aims to tackle scientific tasks directly by predicting target values with neural networks (e.g., physics-informed neural networks, neural ODEs, neural operators, etc.), but attaining high accuracy…

Large Language Models (LLMs) have achieved high accuracy on complex commonsense and mathematical problems that involve the composition of multiple reasoning steps. However, current compositional benchmarks testing these skills tend to focus…

Computation and Language · Computer Science 2026-05-26 Lisa Alazraki , Lihu Chen , Ana Brassard , Joe Stacey , Hossein A. Rahmani , Marek Rei

The GRavitational lEnsing Accuracy Testing 3 (GREAT3) challenge is the third in a series of image analysis challenges, with a goal of testing and facilitating the development of methods for analyzing astronomical images that will be used to…

LLM agents are increasingly deployed as executable systems that use tools, modify workspaces, and produce concrete artifacts. In such workflows, performance depends not only on the base model, but also on the harness: the system layer that…

Artificial Intelligence · Computer Science 2026-05-28 Yilun Yao , Xinyu Tan , Chao-Hsuan Liu , Yaoming Li , Zhengyang Wang , Wenhan Yu , Zhewen Tan , Yuxuan Tian , Guangxiang Zhao , Lin Sun , Xiangzheng Zhang , Tong Yang

We introduce DA-Code, a code generation benchmark specifically designed to assess LLMs on agent-based data science tasks. This benchmark features three core elements: First, the tasks within DA-Code are inherently challenging, setting them…

Computation and Language · Computer Science 2024-10-14 Yiming Huang , Jianwen Luo , Yan Yu , Yitong Zhang , Fangyu Lei , Yifan Wei , Shizhu He , Lifu Huang , Xiao Liu , Jun Zhao , Kang Liu

Gravitational-wave astronomy of compact binaries relies on theoretical models of the gravitational-wave signal that is emitted as binaries coalesce. These models do not only need to be accurate, they also have to be fast to evaluate in…

Instrumentation and Methods for Astrophysics · Physics 2020-03-04 Yoshinta Setyawati , Michael Pürrer , Frank Ohme

Deep learning methods have been employed in gravitational-wave astronomy to accelerate the construction of surrogate waveforms for the inspiral of spin-aligned black hole binaries, among other applications. We face the challenge of modeling…

Instrumentation and Methods for Astrophysics · Physics 2023-08-24 Styliani-Christina Fragkouli , Paraskevi Nousi , Nikolaos Passalis , Panagiotis Iosif , Nikolaos Stergioulas , Anastasios Tefas

We discuss the results of using large language models (LLMs) to conduct original scientific research in an unfamiliar subject area during the Fall 2025 semester. Students in a graduate astronomy and astrophysics course were asked to test…

Instrumentation and Methods for Astrophysics · Physics 2026-03-30 Ann Zabludoff , Chen-Yu Chuang , Parker Thomas Johnson , Yichen Liu , Brina Bianca Martinez , Neev Shah , Lucille Steffes , Gabriel Glen Weible

LLM agents have incredible potential for scientific discovery applications. However, the performance of LLM agents on real-world, small molecule drug design (SMDD) tasks across diverse chemistries and targets is unclear. Current evaluation…

Artificial Intelligence · Computer Science 2026-05-26 Kevin Han , Renfei Zhang , Kathy Wei , Hamed Mahdavi , Niloofar Mireshghallah , Amir Barati Farimani

Proof engineering is notoriously labor-intensive: proofs that are straightforward on paper often require lengthy scripts in theorem provers. Recent advances in large language models (LLMs) create new opportunities for proof automation:…

Programming Languages · Computer Science 2026-01-08 Yichen Xu , Martin Odersky
‹ Prev 1 8 9 10 Next ›