English
Related papers

Related papers: GAIA: a benchmark for General AI Assistants

200 papers

Generative AI, the most popular current approach to AI, consists of large language models (LLMs) that are trained to produce outputs that are plausible, but not necessarily correct. Although their abilities are often uncanny, they are…

Machine Learning · Computer Science 2023-08-10 Doug Lenat , Gary Marcus

Data wrangling tasks such as obtaining and linking data from various sources, transforming data formats, and correcting erroneous records, can constitute up to 80% of typical data engineering work. Despite the rise of machine learning and…

Generative artificial intelligence (GenAI) holds the potential to transform the delivery, cultivation, and evaluation of human learning. This Perspective examines the integration of GenAI as a tool for human learning, addressing its…

Human-Computer Interaction · Computer Science 2024-09-06 Lixiang Yan , Samuel Greiff , Ziwen Teuber , Dragan Gašević

A crucial consideration when developing and deploying Large Language Models (LLMs) is the human values to which these models are aligned. In the constitutional framework of alignment models are aligned to a set of principles (the…

Machine Learning · Computer Science 2026-01-27 Henry Bell , Lara Neubauer da Costa Schertel , Bochu Ding , Brandon Fain

Large language models (LLMs) are advanced artificial intelligence (AI) systems that can perform a variety of tasks commonly found in human intelligence tests, such as defining words, performing calculations, and engaging in verbal…

Computation and Language · Computer Science 2024-09-12 David Ilić , Gilles E. Gignac

Artificial Intelligence Impact Assessments ("AIIAs"), a family of tools that provide structured processes to imagine the possible impacts of a proposed AI system, have become an increasingly popular proposal to govern AI systems. Recent…

Computers and Society · Computer Science 2023-11-21 Nari Johnson , Hoda Heidari

Artificial Intelligence (AI) is rapidly embedded in critical decision-making systems, however their foundational ``black-box'' models require eXplainable AI (XAI) solutions to enhance transparency, which are mostly oriented to experts,…

Machine Learning · Computer Science 2025-06-17 Eva Paraschou , Ioannis Arapakis , Sofia Yfantidou , Sebastian Macaluso , Athena Vakali

Is an LLM telling you different facts than it's telling me? This paper introduces ConsistencyAI, an independent benchmark for measuring the factual consistency of large language models (LLMs) for different personas. ConsistencyAI tests…

Computation and Language · Computer Science 2025-10-30 Peter Banyas , Shristi Sharma , Alistair Simmons , Atharva Vispute

With the rise of artificial intelligence (A.I.) and large language models like ChatGPT, a new race for achieving artificial general intelligence (A.G.I) has started. While many speculate how and when A.I. will achieve A.G.I., there is no…

Artificial Intelligence · Computer Science 2025-12-05 Georgios Mappouras

Large Language Models (LLMs) are revolutionizing the development of AI assistants capable of performing diverse tasks across domains. However, current state-of-the-art LLM-driven agents face significant challenges, including high…

Agent benchmarks remain largely English-centric, while their multilingual versions are often built with machine translation (MT) and limited post-editing. We argue that, for agentic tasks, this minimal workflow can easily break benchmark…

Computation and Language · Computer Science 2026-04-29 Yunsu Kim , Kaden Uhlig , Joern Wuebker

Contemporary large language models (LLMs) have demonstrated remarkable reasoning capabilities, particularly in specialized domains like mathematics and physics. However, their ability to generalize these reasoning skills to more general and…

Computation and Language · Computer Science 2026-04-14 Junlin Liu , Shengnan An , Shuang Zhou , Dan Ma , Shixiong Luo , Ying Xie , Yuan Zhang , Wenling Yuan , Yifan Zhou , Xiaoyu Li , Ziwen Wang , Xuezhi Cao , Xunliang Cai

This technical report describes the AIA Forecaster, a Large Language Model (LLM)-based system for judgmental forecasting using unstructured data. The AIA Forecaster approach combines three core elements: agentic search over high-quality…

Artificial intelligence (AI) has made significant strides in recent years, yet it continues to struggle with a fundamental aspect of cognition present in all animals: common sense. Current AI systems, including those designed for complex…

Artificial Intelligence · Computer Science 2025-01-14 Hugo Latapie

We investigate the presence of cognitive biases in three large language models (LLMs): GPT-4o, Gemma 2, and Llama 3.1. The study uses 1,500 experiments across nine established cognitive biases to evaluate the models' responses and…

Artificial Intelligence · Computer Science 2025-09-11 Payam Saeedi , Mahsa Goodarzi , M Abdullah Canbaz

Generative AI (GenAI), especially Large Language Models (LLMs), is rapidly reshaping both programming workflows and computer science education. Many programmers now incorporate GenAI tools into their workflows, including for collaborative…

Human-Computer Interaction · Computer Science 2025-05-14 Wenhan Lyu , Yimeng Wang , Yifan Sun , Yixuan Zhang

Although general-purpose artificial intelligence (GPAI) is widely expected to accelerate scientific discovery, its practical limits in biomedicine remain unclear. We assess this potential by developing a framework of GPAI capabilities…

Computers and Society · Computer Science 2025-08-26 Konstantin Hebenstreit , Constantin Convalexius , Stephan Reichl , Stefan Huber , Christoph Bock , Matthias Samwald

Deployed artificial intelligence (AI) often impacts humans, and there is no one-size-fits-all metric to evaluate these tools. Human-centered evaluation of AI-based systems combines quantitative and qualitative analysis and human input. It…

Human-Computer Interaction · Computer Science 2023-03-14 Teresa Datta , John P. Dickerson

This study aimed to examine an assumption that generative artificial intelligence (GAI) tools can overcome the cognitive intensity that humans suffer when solving problems. We compared the performance of ChatGPT and GPT-4 on 2019 NAEP…

Artificial Intelligence · Computer Science 2024-01-31 Xiaoming Zhai , Matthew Nyaaba , Wenchao Ma

Recent advancements in artificial intelligence (AI), particularly in large language models (LLMs) such as OpenAI-o1 and DeepSeek-R1, have demonstrated remarkable capabilities in complex domains such as logical reasoning and experimental…

‹ Prev 1 3 4 5 6 7 10 Next ›