中文
相关论文

相关论文: Understanding AI Evaluation Patterns: How Differen…

200 篇论文

AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their proprietary search integrations and retrieval-synthesis pipelines, handle emerging facts…

计算与语言 · 计算机科学 2026-05-22 Mirac Suzgun , Emily Shen , Federico Bianchi , Alexander Spangher , Thomas Icard , Daniel E. Ho , Dan Jurafsky , James Zou

Large-scale AI models such as GPT-4 have accelerated the deployment of artificial intelligence across critical domains including law, healthcare, and finance, raising urgent questions about trust and transparency. This study investigates…

人工智能 · 计算机科学 2025-10-20 Allen Daniel Sunny

Despite recent advances in text-to-3D generative methods, there is a notable absence of reliable evaluation metrics. Existing metrics usually focus on a single criterion each, such as how well the asset aligned with the input text. These…

计算机视觉与模式识别 · 计算机科学 2024-01-11 Tong Wu , Guandao Yang , Zhibing Li , Kai Zhang , Ziwei Liu , Leonidas Guibas , Dahua Lin , Gordon Wetzstein

With the rise of foundation models, a new artificial intelligence paradigm has emerged, by simply using general purpose foundation models with prompting to solve problems instead of training a separate machine learning model for each…

人工智能 · 计算机科学 2023-08-29 Mostafa M. Amin , Rui Mao , Erik Cambria , Björn W. Schuller

Multimodal GPTs represent a watershed in the interplay between Software Engineering and Generative Artificial Intelligence. GPT-4 accepts image and text inputs, rather than simply natural language. We investigate relevant use cases stemming…

软件工程 · 计算机科学 2025-08-21 Roberto Rossi

Generative AI (GenAI) systems are inherently non-deterministic, producing varied outputs even for identical inputs. While this variability is central to their appeal, it challenges established HCI evaluation practices that typically assume…

人机交互 · 计算机科学 2026-01-30 Hyerim Park , Khanh Huynh , Malin Eiband , Jeremy Dillmann , Sven Mayer , Michael Sedlmair

Natural language explanation (NLE) models aim at explaining the decision-making process of a black box system via generating natural language sentences which are human-friendly, high-level and fine-grained. Current NLE models explain the…

计算机视觉与模式识别 · 计算机科学 2022-03-11 Fawaz Sammani , Tanmoy Mukherjee , Nikos Deligiannis

Human cognitive biases in software engineering can lead to costly errors. While general-purpose AI (GPAI) systems may help mitigate these biases due to their non-human nature, their training on human-generated data raises a critical…

The emergence of large language models such as ChatGPT, Gemini, and others highlights the importance of evaluating their diverse capabilities, ranging from natural language understanding to code generation. However, their performance on…

计算与语言 · 计算机科学 2025-01-06 Liuchang Xu , Shuo Zhao , Qingming Lin , Luyao Chen , Qianqian Luo , Sensen Wu , Xinyue Ye , Hailin Feng , Zhenhong Du

The increasing frequency and sophistication of cybersecurity vulnerabilities in software systems underscores the need for more robust and effective vulnerability assessment methods. However, existing approaches often rely on highly…

密码学与安全 · 计算机科学 2025-05-21 Shivansh Chopra , Hussain Ahmad , Diksha Goel , Claudia Szabo

Artificial intelligence (AI) is poised to transform education, but the research community lacks a robust, general benchmark to evaluate AI models for learning. To assess state-of-the-art support for educational use cases, we ran an "arena…

GPT-4 is often heralded as a leading commercial AI offering, sparking debates over its potential as a steppingstone toward artificial general intelligence. But does it possess consciousness? This paper investigates this key question using…

人工智能 · 计算机科学 2024-07-16 Izak Tait , Joshua Bensemann , Ziqi Wang

As Large Multimodal Models (LMMs) become integral to daily digital life, understanding their safety architectures is a critical problem for AI Alignment. This paper presents a systematic analysis of OpenAI's GPT-4o mini, a globally deployed…

机器学习 · 计算机科学 2026-05-26 Niruthiha Selvanayagam , Ted Kurti

We assess whether AI systems can credibly evaluate investment risk appetite-a task that must be thoroughly validated before automation. Our analysis was conducted on proprietary systems (GPT, Claude, Gemini) and open-weight models (LLaMA,…

We introduce VisualQuest, a novel dataset designed to rigorously evaluate multimodal large language models (MLLMs) on abstract visual reasoning tasks that require the integration of symbolic, cultural, and linguistic knowledge. Unlike…

计算机视觉与模式识别 · 计算机科学 2026-01-05 Kelaiti Xiao , Liang Yang , Dongyu Zhang , Paerhati Tulajiang , Hongfei Lin

Large Language Models (LLMs), representing a significant achievement in artificial intelligence (AI) research, have demonstrated their ability in a multitude of tasks. This project aims to explore the capabilities of GPT-3.5, a leading…

计算与语言 · 计算机科学 2023-11-02 Jingjing Wang , Joshua Luo , Grace Yang , Allen Hong , Feng Luo

Recent advances in generative artificial intelligence (AI) have shown promise in accurately grading open-ended student responses. However, few prior works have explored grading handwritten responses due to a lack of data and the challenge…

计算机与社会 · 计算机科学 2024-12-13 Adriana Caraeni , Alexander Scarlatos , Andrew Lan

The increasing reliance on AI-driven solutions, particularly Large Language Models (LLMs) like the GPT series, for information retrieval highlights the critical need for their factuality and fairness, especially amidst the rampant spread of…

计算与语言 · 计算机科学 2024-02-01 Shujaat Mirza , Bruno Coelho , Yuyuan Cui , Christina Pöpper , Damon McCoy

This paper investigates bias in GLLM annotations by conceptually replicating manual annotations of Boukes (2024). Using various GLLMs (Llama3.1:8b, Llama3.3:70b, GPT4o, Qwen2.5:72b) in combination with five different prompts for five…

计算与语言 · 计算机科学 2025-12-10 Sjoerd B. Stolwijk , Mark Boukes , Damian Trilling

The development in Artificial Intelligence (AI) offers transformative potential for redefining student assessment methodologies. This paper aims to establish the idea of the advancement of Artificial Intelligence (AI) and its prospect in…

计算机与社会 · 计算机科学 2025-03-10 Pushpalatha K S , Abhishek Mangalur , Ketan Hegde , Chetan Badachi , Mohammad Aamir