中文
相关论文

相关论文: Benchmark It Yourself (BIY): Preparing a Dataset a…

200 篇论文

Generative AI systems achieve impressive performance on standard benchmarks yet fail to deliver real-world utility, a disconnect we identify across 28 deployment cases spanning education, healthcare, software engineering, and law. We argue…

机器学习 · 计算机科学 2026-05-12 Ishani Mondal , Shweta Bhardwaj

Advances in artificial intelligence need to become more resource-aware and sustainable. This requires clear assessment and reporting of energy efficiency trade-offs, like sacrificing fast running time for higher predictive performance.…

机器学习 · 计算机科学 2023-04-18 Raphael Fischer , Matthias Jakobs , Katharina Morik

We present a novel approach for constructing discrete optimization benchmarks that enables fine-grained control over problem properties, and such benchmarks can facilitate analyzing discrete algorithm behaviors. We build benchmark problems…

神经与进化计算 · 计算机科学 2026-04-09 Furong Ye , Frank Neumann , Thomas Bäck , Niki van Stein

Large language models are proliferating, and so are the benchmarks that serve as their common yardsticks. We ask how the agglomeration patterns of these two layers compare: do they evolve in tandem or diverge? Drawing on two curated proxies…

计算机与社会 · 计算机科学 2025-10-03 Manuel Cebrian , Tomomi Kito , Raul Castro Fernandez

The increasing attention on deep learning has tremendously spurred the design of intelligence processing hardware. The variety of emerging intelligence processors requires standard benchmarks for fair comparison and system optimization (in…

How do multimodal models solve visual spatial tasks -- through genuine planning, or through brute-force search in token space? We introduce \textsc{MazeBench}, a benchmark of 110 procedurally generated maze images across nine controlled…

机器学习 · 计算机科学 2026-05-14 Alberto G. Rodriguez Salgado

Evaluating models on large benchmarks is very resource-intensive, especially during the period of rapid model evolution. Existing efficient evaluation methods estimate the performance of target models by testing them only on a small and…

机器学习 · 计算机科学 2025-06-03 Peiwen Yuan , Yueqi Zhang , Shaoxiong Feng , Yiwei Li , Xinglin Wang , Jiayi Shi , Chuyi Tan , Boyuan Pan , Yao Hu , Kan Li

AI model documentation is fragmented across platforms and inconsistent in structure, preventing policymakers, auditors, and users from reliably assessing safety claims, data provenance, and version-level changes. We analyzed documentation…

人工智能 · 计算机科学 2025-12-16 Akhmadillo Mamirov , Faiaz Azmain , Hanyu Wang

Due to increasing amounts of data and compute resources, deep learning achieves many successes in various domains. The application of deep learning on the mobile and embedded devices is taken more and more attentions, benchmarking and…

机器学习 · 计算机科学 2020-05-12 Chunjie Luo , Xiwen He , Jianfeng Zhan , Lei Wang , Wanling Gao , Jiahui Dai

Frequentist statistical methods, such as hypothesis testing, are standard practice in papers that provide benchmark comparisons. Unfortunately, these methods have often been misused, e.g., without testing for their statistical test…

统计方法学 · 统计学 2021-05-18 David Issa Mattos , Jan Bosch , Helena Holmström Olsson

Big data benchmark suites must include a diversity of data and workloads to be useful in fairly evaluating big data systems and architectures. However, using truly comprehensive benchmarks poses great challenges for the architecture…

性能 · 计算机科学 2016-11-15 Zhen Jia , Jianfeng Zhan , Lei Wang , Rui Han , Sally A. McKee , Qiang Yang , Chunjie Luo , Jingwei Li

The rapid growth of AI agent ecosystems is transforming how complex tasks are delegated and executed, creating a new challenge of identifying suitable agents for a given task. Unlike traditional tools, agent capabilities are often…

人工智能 · 计算机科学 2026-04-27 Bin Wu , Arastun Mammadli , Xiaoyu Zhang , Emine Yilmaz

The meteoric rise of AI, with its rapidly expanding market capitalization, presents both transformative opportunities and critical challenges. Chief among these is the urgent need for a new, unified paradigm for trustworthy evaluation, as…

Numerous benchmarks for Few-Shot Learning have been proposed in the last decade. However all of these benchmarks focus on performance averaged over many tasks, and the question of how to reliably evaluate and tune models trained for…

机器学习 · 计算机科学 2023-07-07 Luísa Shimabucoro , Timothy Hospedales , Henry Gouk

Recent advances in large language models (LLMs) have fueled growing interest in automating geospatial analysis and GIS workflows, yet their actual capabilities remain uncertain. In this work, we call for rigorous evaluation of LLMs on…

软件工程 · 计算机科学 2025-09-09 Qianheng Zhang , Song Gao , Chen Wei , Yibo Zhao , Ying Nie , Ziru Chen , Shijie Chen , Yu Su , Huan Sun

Benchmark datasets play a central role in the organization of machine learning research. They coordinate researchers around shared research problems and serve as a measure of progress towards shared goals. Despite the foundational role of…

机器学习 · 计算机科学 2021-12-06 Bernard Koch , Emily Denton , Alex Hanna , Jacob G. Foster

Recent advances in large language models have enabled the emergence of AI scientists that aim to autonomously analyze biological data and assist scientific discovery. Despite rapid progress, it remains unclear to what extent these systems…

人工智能 · 计算机科学 2026-01-21 Erpai Luo , Jinmeng Jia , Yifan Xiong , Xiangyu Li , Xiaobo Guo , Baoqi Yu , Minsheng Hao , Lei Wei , Xuegong Zhang

In visual interactive labeling, users iteratively assign labels to data items until the machine model reaches an acceptable accuracy. A crucial step of this process is to inspect the model's accuracy and decide whether it is necessary to…

人机交互 · 计算机科学 2021-10-15 Nicolas Grossmann , Jürgen Bernard , Michael Sedlmair , Manuela Waldner

Recent failures such as Google Gemini generating people of color in Nazi-era uniforms illustrate how AI outputs can be factually plausible yet socially harmful. AI models are increasingly evaluated for "fairness," yet existing benchmarks…

计算与语言 · 计算机科学 2025-10-01 Jen-tse Huang , Yuhang Yan , Linqi Liu , Yixin Wan , Wenxuan Wang , Kai-Wei Chang , Michael R. Lyu

This study highlights the potential of fine-tuned ChatGPT (GPT-3.5) for automatically scoring student written constructed responses using example assessment tasks in science education. Recent studies on OpenAI's generative model GPT-3.5…

计算与语言 · 计算机科学 2023-12-27 Ehsan Latif , Xiaoming Zhai