English
Related papers

Related papers: Understanding AI Evaluation Patterns: How Differen…

200 papers

Automated sentiment analysis using Large Language Model (LLM)-based models like ChatGPT, Gemini or LLaMA2 is becoming widespread, both in academic research and in industrial applications. However, assessment and validation of their…

Computation and Language · Computer Science 2024-02-06 Alessio Buscemi , Daniele Proverbio

This study investigates regional bias in large language models (LLMs), an emerging concern in AI fairness and global representation. We evaluate ten prominent LLMs: GPT-3.5, GPT-4o, Gemini 1.5 Flash, Gemini 1.0 Pro, Claude 3 Opus, Claude…

Computation and Language · Computer Science 2026-01-26 M P V S Gopinadh , Kappara Lakshmi Sindhu , Soma Sekhar Pandu Ranga Raju P , Yesaswini Swarna

Large language models (LLMs) such as ChatGPT are increasingly integrated into high-stakes decision-making, yet little is known about their susceptibility to social influence. We conducted three preregistered conformity experiments with…

Artificial Intelligence · Computer Science 2025-10-31 Clarissa Sabrina Arlinghaus , Tristan Kenneweg , Barbara Hammer , Günter W. Maier

Recently, Large Language Models like ChatGPT have demonstrated remarkable proficiency in various Natural Language Processing tasks. Their application in Requirements Engineering, especially in requirements classification, has gained…

Artificial Intelligence · Computer Science 2024-04-17 Abdelkarim El-Hajjami , Nicolas Fafin , Camille Salinesi

GPT-Vision has impressed us on a range of vision-language tasks, but it comes with the familiar new challenge: we have little idea of its capabilities and limitations. In our study, we formalize a process that many have instinctively been…

Computation and Language · Computer Science 2023-11-06 Alyssa Hwang , Andrew Head , Chris Callison-Burch

It has been suggested that large language models such as GPT-4 have acquired some form of understanding beyond the correlations among the words in text including some understanding of mathematics as well. Here, we perform a critical inquiry…

Machine Learning · Computer Science 2023-11-15 Roozbeh Yousefzadeh , Xuenan Cao

Prior work has shown that large language models (LLMs) can predict human attitudes based on other attitudes, but this work has largely focused on predictions from highly similar and interrelated attitudes. In contrast, human attitudes are…

Computation and Language · Computer Science 2025-03-28 Ana Ma , Derek Powell

This paper presents research on enhancements to Large Language Models (LLMs) through the addition of diversity in its generated outputs. Our study introduces a configuration of multiple LLMs which demonstrates the diversities capable with a…

Computation and Language · Computer Science 2025-08-05 Purva Prasad Gosavi , Vaishnavi Murlidhar Kulkarni , Alan F. Smeaton

As Large Language Models (LLMs) have become integral to both research and daily operations, rigorous evaluation is crucial. This assessment is important not only for individual tasks but also for understanding their societal impact and…

Software Engineering · Computer Science 2024-04-02 Zeeshan Rasheed , Muhammad Waseem , Kari Systä , Pekka Abrahamsson

As we consider entrusting Large Language Models (LLMs) with key societal and decision-making roles, measuring their alignment with human cognition becomes critical. This requires methods that can assess how these systems represent…

Artificial Intelligence · Computer Science 2025-10-03 Mattson Ogg , Ritwik Bose , Jamie Scharf , Christopher Ratto , Michael Wolmetz

Many important decisions in our everyday lives, such as authentication via biometric models, are made by Artificial Intelligence (AI) systems. These can be in poor alignment with human expectations, and testing them on clear-cut existing…

Human-Computer Interaction · Computer Science 2024-09-20 Lukas Mecke , Daniel Buschek , Uwe Gruenefeld , Florian Alt

Warning: This paper contains content that may be offensive or upsetting. There has been a significant increase in the usage of large language models (LLMs) in various applications, both in their original form and through fine-tuned…

Computation and Language · Computer Science 2023-12-12 Jiaxu Zhao , Meng Fang , Shirui Pan , Wenpeng Yin , Mykola Pechenizkiy

Industrial teams often deploy large language model features before stable regression or model selection evaluation exists. We present a reusable evaluation system for AI meeting summaries that combines structured ground-truth (GT)…

Artificial Intelligence · Computer Science 2026-05-14 Philip Zhong , Don Wang , Jason Zhang

Large Language Models (LLMs), which simulate human users, are frequently employed to evaluate chatbots in applications such as tutoring and customer service. Effective evaluation necessitates a high degree of human-like diversity within…

Computation and Language · Computer Science 2024-09-04 Xiaoyu Lin , Xinkai Yu , Ankit Aich , Salvatore Giorgi , Lyle Ungar

Background: Evaluating AI-generated treatment plans is a key challenge as AI expands beyond diagnostics, especially with new reasoning models. This study compares plans from human experts and two AI models (a generalist and a reasoner),…

Artificial Intelligence · Computer Science 2025-07-09 Dipayan Sengupta , Saumya Panda

Generative AI and large language models have the potential to drastically improve the landscape of computing education by automatically generating personalized feedback and content. Recent works have studied the capabilities of these models…

Machine Learning · Computer Science 2023-08-08 Adish Singla

Providing effective feedback is important for student learning in programming problem-solving. In this sense, Large Language Models (LLMs) have emerged as potential tools to automate feedback generation. However, their reliability and…

Software Engineering · Computer Science 2025-03-20 Priscylla Silva , Evandro Costa

This work investigates the political biases and personality traits of ChatGPT, specifically comparing GPT-3.5 to GPT-4. In addition, the ability of the models to emulate political viewpoints (e.g., liberal or conservative positions) is…

Computation and Language · Computer Science 2024-10-29 Erik Weber , Jérôme Rutinowski , Niklas Jost , Markus Pauly

In this work, we introduce Mini-Gemini, a simple and effective framework enhancing multi-modality Vision Language Models (VLMs). Despite the advancements in VLMs facilitating basic visual dialog and reasoning, a performance gap persists…

Computer Vision and Pattern Recognition · Computer Science 2024-03-28 Yanwei Li , Yuechen Zhang , Chengyao Wang , Zhisheng Zhong , Yixin Chen , Ruihang Chu , Shaoteng Liu , Jiaya Jia

Large Language Models such as GPTs (Generative Pre-trained Transformers) exhibit remarkable capabilities across a broad spectrum of applications. Nevertheless, due to their intrinsic complexity, these models present substantial challenges…

Machine Learning · Computer Science 2024-10-17 Ashkan Golgoon , Khashayar Filom , Arjun Ravi Kannan
‹ Prev 1 8 9 10 Next ›