English
Related papers

Related papers: Beyond Benchmarks: How Users Evaluate AI Chat Assi…

200 papers

Conversational AI companions have grown prominent in public discourse, yet scholarly understanding of user experiences remains limited, with existing research organized around evaluative poles of harm and benefit rather than examining what…

Human-Computer Interaction · Computer Science 2026-04-09 Dayeon Eom , Julianne Renner , Sedona Chinn

As large language model (LLM)-enhanced chatbots become increasingly expressive and socially responsive, many users begin forming companionship-like bonds with them. This study investigates how using AI companions relates to psychological…

Human-Computer Interaction · Computer Science 2026-05-06 Yutong Zhang , Dora Zhao , Jeffrey T. Hancock , Robert Kraut , Diyi Yang

Recent advancements in large language models, including GPT-4 and its variants, and Generative AI-assisted coding tools like GitHub Copilot, ChatGPT, and Tabnine, have significantly transformed software development. This paper analyzes how…

Software Engineering · Computer Science 2024-11-05 Vijay Joshi , Iver Band

Sentiment analysis is a well-known natural language processing task that involves identifying the emotional tone or polarity of a given piece of text. With the growth of social media and other online platforms, sentiment analysis has become…

Computation and Language · Computer Science 2023-07-03 Mohammad Belal , James She , Simon Wong

Millions of clinicians use ChatGPT to support clinical care, but evaluations of the most common use cases in model-clinician conversations are limited. We introduce HealthBench Professional, an open benchmark for evaluating large language…

In recent years, we have seen an influx in reliance on AI assistants for information seeking. Given this widespread use and the known challenges AI poses for Black users, recent efforts have emerged to identify key considerations needed to…

Human-Computer Interaction · Computer Science 2025-04-21 Lisa Egede , Ebtesam Al Haque , Gabriella Thompson , Alicia Boyd , Angela D. R. Smith , Brittany Johnson

Recent advancements in AI, particularly in large language models (LLMs) like ChatGPT, Claude, and Gemini, have prompted questions about their proximity to Artificial General Intelligence (AGI). This study compares LLM performance on…

Artificial Intelligence · Computer Science 2024-07-16 Mfon Akpan

Millions of patients are already using large language model (LLM) chatbots for medical advice on a regular basis, raising patient safety concerns. This physician-led red-teaming study compares the safety of four publicly available…

Large Language Models (LLMs) like ChatGPT, Copilot, Gemini, and DeepSeek are transforming software engineering by automating key tasks, including code generation, testing, and debugging. As these models become integral to development…

Software Engineering · Computer Science 2025-08-07 Everton Guimaraes , Nathalia Nascimento , Chandan Shivalingaiah , Asish Nelapati

Large Language Models (LLMs) have emerged as powerful tools for tackling a wide range of problems, including those in scientific computing, particularly in solving partial differential equations (PDEs). However, different models exhibit…

Machine Learning · Computer Science 2025-03-11 Qile Jiang , Zhiwei Gao , George Em Karniadakis

Artificial Intelligence (AI) tools such as GitHub Copilot and ChatGPT are increasingly used in software engineering (SE) for tasks such as code, test, and documentation generation. However, engineers often face uncertainty about when to…

Software Engineering · Computer Science 2026-03-24 Vahid Garousi , Zafar Jafarov , Aytan Mövsümova , Atif Namazov

Chatbots have long been capable of answering basic questions and even responding to obscure prompts, but recently their improvements have been far more significant. Modern chatbots like Open AIs ChatGPT3 not only have the ability to answer…

Computation and Language · Computer Science 2023-03-22 Grant Rosario , David Noever

Evaluating conversational AI systems that use external tools is challenging, as errors can arise from complex interactions among user, agent, and tools. While existing evaluation methods assess either user satisfaction or agents'…

Computation and Language · Computer Science 2025-10-23 Zhaoyi Joey Hou , Tanya Shourya , Yingfan Wang , Shamik Roy , Vinayshekhar Bannihatti Kumar , Rashmi Gangadharaiah

Large language models have gained considerable interest for their impressive performance on various tasks. Among these models, ChatGPT developed by OpenAI has become extremely popular among early adopters who even regard it as a disruptive…

Computation and Language · Computer Science 2023-04-10 Aman Rangapur , Haoran Wang

ChatGPT, powered by a large language model (LLM), has revolutionized everyday human-computer interaction (HCI) since its 2022 release. While now used by millions around the world, a coherent pathway for evaluating the user experience (UX)…

Human-Computer Interaction · Computer Science 2025-03-21 Katie Seaborn

Many people suffer from mental health problems but not everyone seeks professional help or has access to mental health care. AI chatbots have increasingly become a go-to for individuals who either have mental disorders or simply want…

Other Quantitative Biology · Quantitative Biology 2025-11-03 Aditya Naik , Jovi Thomas , Teja Sree Mandava , Himavanth Reddy Vemula

The rapid evolution of LLMs represents an impactful paradigm shift in digital interaction and content engagement. While they encode vast amounts of human-generated knowledge and excel in processing diverse data types, they often face the…

Human-Computer Interaction · Computer Science 2024-11-20 Anna Bodonhelyi , Efe Bozkir , Shuo Yang , Enkelejda Kasneci , Gjergji Kasneci

User satisfaction with AI assistants is highly personalized: the same response may satisfy one user but disappoint another depending on what each user expects and what they have asked for before. Existing automatic evaluation methods mostly…

Computation and Language · Computer Science 2026-05-29 Zhefan Wang , Zhiqiang Guo , Weizhi Ma , Min Zhang , Quanjia Yan , Hengliang Luo

This study compares the performance of AI-generated and human-written product descriptions using a multifaceted evaluation model. We analyze descriptions for 100 products generated by four AI models (Gemma 2B, LLAMA, GPT2, and ChatGPT 4)…

Computation and Language · Computer Science 2024-12-30 Sanjukta Ghosh

A rapidly growing body of research is examining how LLMs influence developers when they code. To date, this research has tended to focus on productivity and code quality outcomes, rather than the underlying cognitive processes involved in…

‹ Prev 1 4 5 6 7 8 10 Next ›