English
Related papers

Related papers: EvalAI: Towards Better Evaluation Systems for AI A…

200 papers

Auto-scaling is an automated approach that dynamically provisions resources for microservices to accommodate fluctuating workloads. Despite the introduction of many sophisticated auto-scaling algorithms, evaluating auto-scalers remains…

Software Engineering · Computer Science 2025-04-14 Shuaiyu Xie , Jian Wang , Yang Luo , Yunqing Yong , Yuzhen Tan , Bing Li

AI systems are becoming increasingly complex, ubiquitous and autonomous, leading to increasing concerns about their impacts on individuals and society. In response, researchers have begun investigating how to ensure that the methods…

Multiagent Systems · Computer Science 2026-04-09 Stephen Cranefield , Nir Oren

The integration of Artificial Intelligence in the development of computer systems presents a new challenge: make intelligent systems explainable to humans. This is especially vital in the field of health and well-being, where transparency…

Machine learning (ML) and artificial intelligence (AI) have become hot topics in many information processing areas, from chatbots to scientific data analysis. At the same time, there is uncertainty about the possibility of extending…

Artificial Intelligence · Computer Science 2018-06-08 Abel Torres Montoya

Artificial Intelligence (AI) and Machine Learning (ML) have significantly impacted various industries, including software development. Software testing, a crucial part of the software development lifecycle (SDLC), ensures the quality and…

Software Engineering · Computer Science 2024-09-05 Ahmed Ramadan , Husam Yasin , Burhan Pektas

We present ParlAI Vote, an interactive web platform for exploring European Parliament debates and votes, and for testing LLMs on vote prediction and bias analysis. This web system connects debate topics, speeches, and roll-call outcomes,…

Computation and Language · Computer Science 2025-12-03 Wenjie Lin , Hange Liu , Yingying Zhuang , Xutao Mao , Jingwei Shi , Xudong Han , Tianyu Shi , Jinrui Yang

Recently, there has been a national push to use machine learning (ML) and artificial intelligence (AI) to advance engineering techniques in all disciplines ranging from advanced fracture mechanics in materials science to soil and water…

Computers and Society · Computer Science 2023-04-25 Andrew Schulz , Suzanne Stathatos , Cassandra Shriver , Roxanne Moore

Both in the domains of Feature Selection and Interpretable AI, there exists a desire to `rank' features based on their importance. Such feature importance rankings can then be used to either: (1) reduce the dataset size or (2) interpret the…

Machine Learning · Computer Science 2022-07-12 Jeroen G. S. Overschie

The advent of advanced AI underscores the urgent need for comprehensive safety evaluations, necessitating collaboration across communities (i.e., AI, software engineering, and governance). However, divergent practices and terminologies…

Software Engineering · Computer Science 2024-05-17 Boming Xia , Qinghua Lu , Liming Zhu , Zhenchang Xing

Many ML models are opaque to humans, producing decisions too complex for humans to easily understand. In response, explainable artificial intelligence (XAI) tools that analyze the inner workings of a model have been created. Despite these…

Computers and Society · Computer Science 2021-06-17 Kiana Alikhademi , Brianna Richardson , Emma Drobina , Juan E. Gilbert

Explainable Artificial Intelligence (XAI) is central to the debate on integrating Artificial Intelligence (AI) and Machine Learning (ML) algorithms into clinical practice. High-performing AI/ML models, such as ensemble learners and deep…

Machine Learning · Computer Science 2024-07-30 Alessandro De Carlo , Enea Parimbelli , Nicola Melillo , Giovanna Nicora

Vision-Language Navigation (VLN) aims to guide agents by leveraging language instructions and visual cues, playing a pivotal role in embodied AI. Indoor VLN has been extensively studied, whereas outdoor aerial VLN remains underexplored. The…

Benchmark-based evaluation remains important for tracking frontier AI progress. But it can both overstate and understate deployed capability because it privileges tasks that can be precisely specified, automatically graded, easy to optimize…

While research on explainable AI (XAI) is booming and explanation techniques have proven promising in many application domains, standardised human-centred evaluation procedures are still missing. In addition, current evaluation procedures…

Human-Computer Interaction · Computer Science 2025-06-18 Ivania Donoso-Guzmán , Jeroen Ooge , Denis Parra , Katrien Verbert

Markets are a promising way to coordinate AI agent activity for similar reasons to those used to justify markets more broadly. In order to effectively participate in markets, agents need to have informative signals of their own ability to…

Artificial Intelligence · Computer Science 2026-04-28 Andrey Fradkin , Rohit Krishnan

Evaluating Large Language Models (LLMs) as general-purpose agents is essential for understanding their capabilities and facilitating their integration into practical applications. However, the evaluation process presents substantial…

Computation and Language · Computer Science 2024-12-25 Chang Ma , Junlei Zhang , Zhihao Zhu , Cheng Yang , Yujiu Yang , Yaohui Jin , Zhenzhong Lan , Lingpeng Kong , Junxian He

This paper introduces a novel framework that accelerates the discovery of actionable relationships in high-dimensional temporal data by integrating machine learning (ML), explainable AI (XAI), and natural language processing (NLP) to…

Machine Learning · Computer Science 2025-06-09 Jiztom Kavalakkatt Francis , Matthew J Darr

A common method to solve complex problems in software engineering, is to divide the problem into multiple sub-problems. Inspired by this, we propose a Modular Architecture for Software-engineering AI (MASAI) agents, where different…

Artificial Intelligence · Computer Science 2024-06-18 Daman Arora , Atharv Sonwane , Nalin Wadhwa , Abhav Mehrotra , Saiteja Utpala , Ramakrishna Bairi , Aditya Kanade , Nagarajan Natarajan

We present CAISAR, an open-source platform under active development for the characterization of AI systems' robustness and safety. CAISAR provides a unified entry point for defining verification problems by using WhyML, the mature and…

Artificial Intelligence · Computer Science 2022-06-22 Julien Girard-Satabin , Michele Alberti , François Bobot , Zakaria Chihani , Augustin Lemesle

The emergence and continued reliance on the Internet and related technologies has resulted in the generation of large amounts of data that can be made available for analyses. However, humans do not possess the cognitive capabilities to…

Machine Learning · Computer Science 2021-01-12 MohammadNoor Injadat , Abdallah Moubayed , Ali Bou Nassif , Abdallah Shami
‹ Prev 1 8 9 10 Next ›