English
Related papers

Related papers: Teaching AI Through Benchmark Construction: QuestB…

200 papers

Commonsense question-answering (QA) tasks, in the form of benchmarks, are constantly being introduced for challenging and comparing commonsense QA systems. The benchmarks provide question sets that systems' developers can use to train and…

Artificial Intelligence · Computer Science 2020-12-23 Henrique Santos , Minor Gordon , Zhicheng Liang , Gretchen Forbush , Deborah L. McGuinness

Cyberharassment is a critical, socially relevant cybersecurity problem because of the adverse effects it can have on targeted groups or individuals. While progress has been made in understanding cyber-harassment, its detection, attacks on…

Computers and Society · Computer Science 2024-05-17 Ebuka Okpala , Nishant Vishwamitra , Keyan Guo , Song Liao , Long Cheng , Hongxin Hu , Yongkai Wu , Xiaohong Yuan , Jeannette Wade , Sajad Khorsandroo

Every AI benchmark operationalizes theoretical assumptions about the capability it claims to assess. When assumptions function as unexamined commitments, benchmarks stabilize the dominant paradigm by narrowing what counts as progress. Over…

Artificial Intelligence · Computer Science 2026-05-15 Theodore J Kalaitzidis

Large Language Models are commonly judged by their scores on standard benchmarks, yet such scores often overstate real capability since they mask the mix of skills a task actually demands. For example, ARC is assumed to test reasoning,…

Computation and Language · Computer Science 2025-10-03 Dongjun Kim , Gyuho Shim , Yongchan Chun , Minhyuk Kim , Chanjun Park , Heuiseok Lim

The recent shift in Generative AI (GenAI) applications from cloud-only environments to end-user devices introduces new challenges in resource management, system efficiency, and user experience. This paper presents ConsumerBench, a…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-06-24 Yile Gu , Rohan Kadekodi , Hoang Nguyen , Keisuke Kamahori , Yiyu Liu , Baris Kasikci

Recent years witness a trend of applying large-scale distributed deep learning algorithms (HPC AI) in both business and scientific computing areas, whose goal is to speed up the training time to achieve a state-of-the-art quality. The HPC…

Performance · Computer Science 2021-02-26 Zihan Jiang , Wanling Gao , Fei Tang , Xingwang Xiong , Lei Wang , Chuanxin Lan , Chunjie Luo , Hongxiao Li , Jianfeng Zhan

The rise of Generative AI (GenAI) tools like ChatGPT has created new opportunities and challenges for computing education. Existing research has primarily focused on GenAI's ability to complete educational tasks and its impact on student…

Software Engineering · Computer Science 2025-11-18 Rufeng Chen , Shuaishuai Jiang , Jiyun Shen , AJung Moon , Lili Wei

Fetching, which includes approaching, grasping, and retrieving, is a critical challenge for robot manipulation tasks. Existing methods primarily focus on table-top scenarios, which do not adequately capture the complexities of environments…

Robotics · Computer Science 2024-10-21 Beining Han , Meenal Parakh , Derek Geng , Jack A Defay , Gan Luyang , Jia Deng

The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified test-time contextual cues, such as hypothetical scenarios, as a source of verbalized…

Computation and Language · Computer Science 2026-05-28 Katharina Deckenbach , Haritz Puerto , Jonas Geiping , Sahar Abdelnabi

Creating fair AI systems is a complex problem that involves the assessment of context-dependent bias concerns. Existing research and programming libraries express specific concerns as measures of bias that they aim to constrain or mitigate.…

Machine Learning · Computer Science 2024-05-30 Emmanouil Krasanakis , Symeon Papadopoulos

Large Language Models (LLMs) have made significant strides in front-end code generation. However, existing benchmarks exhibit several critical limitations: many tasks are overly simplistic, test cases often lack rigor, and end-to-end…

Software Engineering · Computer Science 2025-06-19 Hongda Zhu , Yiwen Zhang , Bing Zhao , Jingzhe Ding , Siyao Liu , Tong Liu , Dandan Wang , Yanan Liu , Zhaojian Li

Despite AI tools becoming more prevalent and applicable to a variety of workplaces, workers consistently report uncertainty about where AI applies, what problems it can help solve, and how it fits into real workflows. In other words, there…

Human-Computer Interaction · Computer Science 2026-04-01 Aakanksha Khandwaha , Edith Law

Evaluation of students' performance for the completion of courses has been a major problem for both students and faculties during the work-from-home period in this COVID pandemic situation. To this end, this paper presents an in-depth…

Machine Learning · Computer Science 2020-09-08 Vipul Bansal , Himanshu Buckchash , Balasubramanian Raman

Large Language Model (LLM) agents have shown great potential for solving real-world problems and promise to be a solution for tasks automation in industry. However, more benchmarks are needed to systematically evaluate automation agents…

Artificial Intelligence · Computer Science 2025-07-16 Yinsheng Li , Zhen Dong , Yi Shao

Background: Recently, ChatGPT and similar generative AI models have attracted hundreds of millions of users and become part of the public discourse. Many believe that such models will disrupt society and will result in a significant change…

Computation and Language · Computer Science 2023-04-28 Steffen Herbold , Annette Hautli-Janisz , Ute Heuer , Zlata Kikteva , Alexander Trautsch

Training certifiably robust neural networks is an important but challenging task. While many algorithms for (deterministic) certified training have been proposed, they are often evaluated on different training schedules, certification…

Machine Learning · Computer Science 2025-05-29 Yuhao Mao , Stefan Balauca , Martin Vechev

Generative AI's emphasis on automation and efficiency challenges design education, where learning is grounded in exploration, reflection, and responsibility. This work introduces AI Craftsmanship, a value-oriented framework drawing on…

Human-Computer Interaction · Computer Science 2026-04-09 Tuan-Ting Huang , Janet Yi-Ching Huang , Stephan Wensveen

Formal theorem-proving benchmarks enable mechanically verifiable evaluation of mathematical reasoning in large language models. However, existing benchmarks mainly focus on Olympiad-style problems and algebraic domains, leaving…

Artificial Intelligence · Computer Science 2026-05-19 Wentao Long , Yunfei Zhang , Chenyi Li , Li Zhou , Chumin Sun , Zaiwen Wen

As large language models continue to advance, their application in educational contexts remains underexplored and under-optimized. In this paper, we address this gap by introducing the first diverse benchmark tailored for educational…

Computation and Language · Computer Science 2026-01-07 Bin Xu , Yu Bai , Huashan Sun , Yiguan Lin , Siming Liu , Xinyue Liang , Yaolin Li , Zhuangzhi Dong , Jingren Zhang , Yufan Deng , Xinyu Zou , Yang Gao , Heyan Huang

With the rapid adoption of LLM-based chatbots, there is a pressing need to evaluate what humans and LLMs can achieve together. However, standard benchmarks, such as MMLU, measure LLM capabilities in isolation (i.e., "AI-alone"). Here, we…

Computation and Language · Computer Science 2025-08-13 Serina Chang , Ashton Anderson , Jake M. Hofman
‹ Prev 1 8 9 10 Next ›