English
Related papers

Related papers: Mars-Bench: A Benchmark for Evaluating Foundation …

200 papers

AI models are increasingly prevalent in high-stakes environments, necessitating thorough assessment of their capabilities and risks. Benchmarks are popular for measuring these attributes and for comparing model performance, tracking…

Artificial Intelligence · Computer Science 2024-11-21 Anka Reuel , Amelia Hardy , Chandler Smith , Max Lamparth , Malcolm Hardy , Mykel J. Kochenderfer

There is widespread optimism that frontier Large Language Models (LLMs) and LLM-augmented systems have the potential to rapidly accelerate scientific discovery across disciplines. Today, many benchmarks exist to measure LLM knowledge and…

Astronomical image interpretation presents a significant challenge for applying multimodal large language models (MLLMs) to specialized scientific tasks. Existing benchmarks focus on general multimodal capabilities but fail to capture the…

Instrumentation and Methods for Astrophysics · Physics 2025-10-22 Jinghang Shi , Xiaoyu Tang , Yang Huang , Yuyang Li , Xiao Kong , Yanxia Zhang , Caizhan Yue

LLM development has aroused great interest in Sequential Recommendation (SR) applications. However, comprehensive evaluation of SR models remains lacking due to the limitations of the existing benchmarks: 1) an overemphasis on accuracy,…

Information Retrieval · Computer Science 2026-04-14 Jianhong Li , Zeheng Qian , Wangze Ni , Haoyang Li , Hongwei Yao , Yang Bai , Kui Ren

Benchmarks are the de facto standard for tracking progress in large language models (LLMs), yet static test sets can rapidly saturate, become vulnerable to contamination, and are costly to refresh. Scalable evaluation of open-ended items…

Computation and Language · Computer Science 2026-03-24 Yandan Zheng , Haoran Luo , Zhenghong Lin , Wenjin Liu , Luu Anh Tuan

Foundation Models (FMs), e.g., large language models, possess attributes of intelligence which offer promise to endow a robot with the contextual understanding necessary to navigate complex, unstructured tasks in the wild. We see three core…

Planetary rover systems need to perform terrain segmentation to identify drivable areas as well as identify specific types of soil for sample collection. The latest Martian terrain segmentation methods rely on supervised learning which is…

Computer Vision and Pattern Recognition · Computer Science 2022-02-03 Edwin Goh , Jingdao Chen , Brian Wilson

Mars, lacking an intrinsic dynamo, is an ideal laboratory to comparatively study induced magnetospheres, which can be found in other terrestrial bodies as well as comets. Additionally, Mars is of particular interest to further exploration…

Evaluation of foundation models often rely on aggregate scores from benchmarks that lack comprehensive coverage and metadata for a fine-grained evaluation. We introduce a framework for automated benchmark generation. Our framework generates…

Foundation models hold significant potential for enabling robots to perform long-horizon general manipulation tasks. However, the simplicity of tasks and the uniformity of environments in existing benchmarks restrict their effective…

Robotics · Computer Science 2025-04-04 Liming Zheng , Feng Yan , Fanfan Liu , Chengjian Feng , Zhuoliang Kang , Lin Ma

Accurate parsing of citations is necessary for machine-readable scholarly infrastructure. But, despite sustained interest in this problem, existing evaluation techniques are often not generalizable, based on synthetic data, or not publicly…

Digital Libraries · Computer Science 2026-03-27 Parth Sarin , Juan Pablo Alperin , Adam Buttrick , Dione Mentis

Artificial Intelligence (AI) technologies have profoundly transformed the field of remote sensing, revolutionizing data collection, processing, and analysis. Traditionally reliant on manual interpretation and task-specific models, remote…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Siqi Lu , Junlin Guo , James R Zimmer-Dauphinee , Jordan M Nieusma , Xiao Wang , Parker VanValkenburgh , Steven A Wernke , Yuankai Huo

Large Atomistic Models (LAMs) have undergone remarkable progress recently, emerging as universal or fundamental representations of the potential energy surface defined by the first-principles calculations of atomistic systems. However, our…

Deep Research Systems (DRS) aim to help users search the web, synthesize information, and deliver comprehensive investigative reports. However, how to rigorously evaluate these systems remains under-explored. Existing deep-research…

Computation and Language · Computer Science 2026-02-02 Ruizhe Li , Mingxuan Du , Benfeng Xu , Chiwei Zhu , Xiaorui Wang , Zhendong Mao

Project-Based Learning (PBL) involves a variety of highly correlated multimodal data, making it a vital educational approach within STEM disciplines. With the rapid development of multimodal large language models (MLLMs), researchers have…

Computation and Language · Computer Science 2025-11-04 Xinyi Wu , Yanhao Jia , Qinglin Zhang , Yiran Qin , Luwei Xiao , Shuai Zhao

Recent advances in generative foundational models, often termed "world models," have propelled interest in applying them to critical tasks like robotic planning and autonomous system training. For reliable deployment, these models must…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Rishi Upadhyay , Howard Zhang , Jim Solomon , Ayush Agrawal , Pranay Boreddy , Shruti Satya Narayana , Yunhao Ba , Alex Wong , Celso M de Melo , Achuta Kadambi

Nucleotide sequences constitute the fundamental genetic basis of biological systems, rendering viral genomic analysis critical for biomedical advancement. Despite progress in biological foundation models, specifically nucleotide foundation…

Machine Learning · Computer Science 2026-05-26 Dongxin Ye , Fang Hu , Han Hu , Shu Hu , Yang Tan , Wanli Ouyang , Stan Z. Li , Jie Cui , Nanqing Dong

Inverse design of materials has significantly advanced target-driven formulation optimization, yet existing materials machine learning benchmarks remain limited to forward property prediction, failing to systematically evaluate inverse…

Materials Science · Physics 2026-05-27 Linhan Wu , Chenxi Wang , Chuhan Yang , Zhengwei Yang , Yuyang Liu

Tabular Foundation Models have recently established the state of the art in supervised tabular learning, by leveraging pretraining to learn generalizable representations of numerical and categorical structured data. However, they lack…

Understanding research papers remains challenging for foundation models due to specialized scientific discourse and complex figures and tables, yet existing benchmarks offer limited fine-grained evaluation at scale. To address this gap, we…

Computation and Language · Computer Science 2026-05-01 Yelin Chen , Fanjin Zhang , Suping Sun , Yunhe Pang , Yuanchun Wang , Jian Song , Xiaoyan Li , Lei Hou , Shu Zhao , Jie Tang , Juanzi Li