中文
相关论文

相关论文: Understanding AI Evaluation Patterns: How Differen…

200 篇论文

As ChatGPT goes viral, generative AI (AIGC, a.k.a AI-generated content) has made headlines everywhere because of its ability to analyze and create text, images, and beyond. With such overwhelming media coverage, it is almost impossible for…

Heuristic evaluation is a widely used method in Human-Computer Interaction (HCI) to inspect interfaces and identify issues based on heuristics. Recently, Large Language Models (LLMs), such as GPT-4o, have been applied in HCI to assist in…

Advances in automated scoring are closely aligned with advances in machine-learning and natural-language-processing techniques. With recent progress in large language models (LLMs), the use of ChatGPT, Gemini, Claude, and other…

计算与语言 · 计算机科学 2025-09-30 Haowei Hua , Hong Jiao , Dan Song

More and more people are experiencing pressure from work, life, and education. These pressures often lead to an anxious state of mind, or even the early symptoms of suicidal ideation. With the advancement of artificial intelligence (AI)…

人机交互 · 计算机科学 2025-03-21 Longdi Xian , Junhao Xu

Evaluating the quality of automatically generated image descriptions is challenging, requiring metrics that capture various aspects such as grammaticality, coverage, correctness, and truthfulness. While human evaluation offers valuable…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Jia-Hong Huang , Hongyi Zhu , Yixian Shen , Stevan Rudinac , Alessio M. Pacces , Evangelos Kanoulas

With the rise of multimodal large language models, GPT-4o stands out as a pioneering model, driving us to evaluate its capabilities. This report assesses GPT-4o across various tasks to analyze its audio processing and reasoning abilities.…

计算与语言 · 计算机科学 2025-02-17 Yu-Xiang Lin , Chih-Kai Yang , Wei-Chih Chen , Chen-An Li , Chien-yu Huang , Xuanjun Chen , Hung-yi Lee

This paper examines the comparative effectiveness of a specialized compiled language model and a general-purpose model like OpenAI's GPT-3.5 in detecting SDGs within text data. It presents a critical review of Large Language Models (LLMs),…

计算与语言 · 计算机科学 2023-07-31 Arash Hajikhani , Carolyn Cole

The rapid advancement of native multi-modal models and omni-models, exemplified by GPT-4o, Gemini, and o3, with their capability to process and generate content across modalities such as text and images, marks a significant milestone in the…

Generative AI (genAI) tools, such as ChatGPT or Copilot, are advertised to improve developer productivity and are being integrated into software development. However, misaligned trust, skepticism, and usability concerns can impede the…

Current AI systems minimize risk by enforcing ideological neutrality, yet this may introduce automation bias by suppressing cognitive engagement in human decision-making. We conducted randomized trials with 2,500 participants to test…

人机交互 · 计算机科学 2025-08-21 Shiyang Lai , Junsol Kim , Nadav Kunievsky , Yujin Potter , James Evans

The rapid development of Generative AI is bringing innovative changes to education and assessment. As the prevalence of students utilizing AI for assignments increases, concerns regarding academic integrity and the validity of assessments…

人工智能 · 计算机科学 2025-12-18 Seok-Hyun Ga , Chun-Yen Chang

As Generative AI rises in adoption, its use has expanded to include domains such as hiring and recruiting. However, without examining the potential of bias, this may negatively impact marginalized populations, including people with…

计算机与社会 · 计算机科学 2024-05-24 Kate Glazko , Yusuf Mohammed , Ben Kosa , Venkatesh Potluri , Jennifer Mankoff

The developments in Generative AI technologies have paved the way for numerous innovations in different fields. Recently, Generative AI has been proposed as a competitor to AES systems in evaluating student essays automatically. Considering…

计算与语言 · 计算机科学 2025-10-20 Enis Oğuz

Gender bias in artificial intelligence (AI) and natural language processing has garnered significant attention due to its potential impact on societal perceptions and biases. This research paper aims to analyze gender bias in Large Language…

计算与语言 · 计算机科学 2023-09-04 Vishesh Thakur

This study investigates the consistency of feedback ratings generated by OpenAI's GPT-4, a state-of-the-art artificial intelligence language model, across multiple iterations, time spans and stylistic variations. The model rated responses…

计算与语言 · 计算机科学 2024-01-19 Veronika Hackl , Alexandra Elena Müller , Michael Granitzer , Maximilian Sailer

Recently, GPT-4 with Vision (GPT-4V) has demonstrated remarkable visual capabilities across various tasks, but its performance in emotion recognition has not been fully evaluated. To bridge this gap, we present the quantitative evaluation…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Zheng Lian , Licai Sun , Haiyang Sun , Kang Chen , Zhuofan Wen , Hao Gu , Bin Liu , Jianhua Tao

Cognitive psychology delves on understanding perception, attention, memory, language, problem-solving, decision-making, and reasoning. Large language models (LLMs) are emerging as potent tools increasingly capable of performing human-level…

计算与语言 · 计算机科学 2023-04-13 Sifatkaur Dhingra , Manmeet Singh , Vaisakh SB , Neetiraj Malviya , Sukhpal Singh Gill

As leading examples of large language models, ChatGPT and Gemini claim to provide accurate and unbiased information, emphasizing their commitment to political neutrality and avoidance of personal bias. This research investigates the…

计算与语言 · 计算机科学 2025-04-10 Dogus Yuksel , Mehmet Cem Catalbas , Bora Oc

OpenAI has released the Chat Generative Pre-trained Transformer (ChatGPT) and revolutionized the approach in artificial intelligence to human-model interaction. Several publications on ChatGPT evaluation test its effectiveness on well-known…

This paper provides an in-depth evaluation of three state-of-the-art Large Language Models (LLMs) for personalized career mentoring in the computing field, using three distinct student profiles that consider gender, race, and professional…

计算与语言 · 计算机科学 2024-12-16 Xiao Luo , Sean O'Connell , Shamima Mithun