English
Related papers

Related papers: Understanding AI Evaluation Patterns: How Differen…

200 papers

AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their proprietary search integrations and retrieval-synthesis pipelines, handle emerging facts…

Computation and Language · Computer Science 2026-05-22 Mirac Suzgun , Emily Shen , Federico Bianchi , Alexander Spangher , Thomas Icard , Daniel E. Ho , Dan Jurafsky , James Zou

Large-scale AI models such as GPT-4 have accelerated the deployment of artificial intelligence across critical domains including law, healthcare, and finance, raising urgent questions about trust and transparency. This study investigates…

Artificial Intelligence · Computer Science 2025-10-20 Allen Daniel Sunny

Despite recent advances in text-to-3D generative methods, there is a notable absence of reliable evaluation metrics. Existing metrics usually focus on a single criterion each, such as how well the asset aligned with the input text. These…

Computer Vision and Pattern Recognition · Computer Science 2024-01-11 Tong Wu , Guandao Yang , Zhibing Li , Kai Zhang , Ziwei Liu , Leonidas Guibas , Dahua Lin , Gordon Wetzstein

With the rise of foundation models, a new artificial intelligence paradigm has emerged, by simply using general purpose foundation models with prompting to solve problems instead of training a separate machine learning model for each…

Artificial Intelligence · Computer Science 2023-08-29 Mostafa M. Amin , Rui Mao , Erik Cambria , Björn W. Schuller

Multimodal GPTs represent a watershed in the interplay between Software Engineering and Generative Artificial Intelligence. GPT-4 accepts image and text inputs, rather than simply natural language. We investigate relevant use cases stemming…

Software Engineering · Computer Science 2025-08-21 Roberto Rossi

Generative AI (GenAI) systems are inherently non-deterministic, producing varied outputs even for identical inputs. While this variability is central to their appeal, it challenges established HCI evaluation practices that typically assume…

Human-Computer Interaction · Computer Science 2026-01-30 Hyerim Park , Khanh Huynh , Malin Eiband , Jeremy Dillmann , Sven Mayer , Michael Sedlmair

Natural language explanation (NLE) models aim at explaining the decision-making process of a black box system via generating natural language sentences which are human-friendly, high-level and fine-grained. Current NLE models explain the…

Computer Vision and Pattern Recognition · Computer Science 2022-03-11 Fawaz Sammani , Tanmoy Mukherjee , Nikos Deligiannis

Human cognitive biases in software engineering can lead to costly errors. While general-purpose AI (GPAI) systems may help mitigate these biases due to their non-human nature, their training on human-generated data raises a critical…

Human-Computer Interaction · Computer Science 2025-12-02 Francesco Sovrano , Gabriele Dominici , Rita Sevastjanova , Alessandra Stramiglio , Alberto Bacchelli

The emergence of large language models such as ChatGPT, Gemini, and others highlights the importance of evaluating their diverse capabilities, ranging from natural language understanding to code generation. However, their performance on…

Computation and Language · Computer Science 2025-01-06 Liuchang Xu , Shuo Zhao , Qingming Lin , Luyao Chen , Qianqian Luo , Sensen Wu , Xinyue Ye , Hailin Feng , Zhenhong Du

The increasing frequency and sophistication of cybersecurity vulnerabilities in software systems underscores the need for more robust and effective vulnerability assessment methods. However, existing approaches often rely on highly…

Cryptography and Security · Computer Science 2025-05-21 Shivansh Chopra , Hussain Ahmad , Diksha Goel , Claudia Szabo

Artificial intelligence (AI) is poised to transform education, but the research community lacks a robust, general benchmark to evaluate AI models for learning. To assess state-of-the-art support for educational use cases, we ran an "arena…

GPT-4 is often heralded as a leading commercial AI offering, sparking debates over its potential as a steppingstone toward artificial general intelligence. But does it possess consciousness? This paper investigates this key question using…

Artificial Intelligence · Computer Science 2024-07-16 Izak Tait , Joshua Bensemann , Ziqi Wang

As Large Multimodal Models (LMMs) become integral to daily digital life, understanding their safety architectures is a critical problem for AI Alignment. This paper presents a systematic analysis of OpenAI's GPT-4o mini, a globally deployed…

Machine Learning · Computer Science 2026-05-26 Niruthiha Selvanayagam , Ted Kurti

We assess whether AI systems can credibly evaluate investment risk appetite-a task that must be thoroughly validated before automation. Our analysis was conducted on proprietary systems (GPT, Claude, Gemini) and open-weight models (LLaMA,…

We introduce VisualQuest, a novel dataset designed to rigorously evaluate multimodal large language models (MLLMs) on abstract visual reasoning tasks that require the integration of symbolic, cultural, and linguistic knowledge. Unlike…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Kelaiti Xiao , Liang Yang , Dongyu Zhang , Paerhati Tulajiang , Hongfei Lin

Large Language Models (LLMs), representing a significant achievement in artificial intelligence (AI) research, have demonstrated their ability in a multitude of tasks. This project aims to explore the capabilities of GPT-3.5, a leading…

Computation and Language · Computer Science 2023-11-02 Jingjing Wang , Joshua Luo , Grace Yang , Allen Hong , Feng Luo

Recent advances in generative artificial intelligence (AI) have shown promise in accurately grading open-ended student responses. However, few prior works have explored grading handwritten responses due to a lack of data and the challenge…

Computers and Society · Computer Science 2024-12-13 Adriana Caraeni , Alexander Scarlatos , Andrew Lan

The increasing reliance on AI-driven solutions, particularly Large Language Models (LLMs) like the GPT series, for information retrieval highlights the critical need for their factuality and fairness, especially amidst the rampant spread of…

Computation and Language · Computer Science 2024-02-01 Shujaat Mirza , Bruno Coelho , Yuyuan Cui , Christina Pöpper , Damon McCoy

This paper investigates bias in GLLM annotations by conceptually replicating manual annotations of Boukes (2024). Using various GLLMs (Llama3.1:8b, Llama3.3:70b, GPT4o, Qwen2.5:72b) in combination with five different prompts for five…

Computation and Language · Computer Science 2025-12-10 Sjoerd B. Stolwijk , Mark Boukes , Damian Trilling

The development in Artificial Intelligence (AI) offers transformative potential for redefining student assessment methodologies. This paper aims to establish the idea of the advancement of Artificial Intelligence (AI) and its prospect in…

Computers and Society · Computer Science 2025-03-10 Pushpalatha K S , Abhishek Mangalur , Ketan Hegde , Chetan Badachi , Mohammad Aamir