English
Related papers

Related papers: Evaluating vision-capable chatbots in interpreting…

200 papers

The use of large language model (LLM)-powered chatbots, such as ChatGPT, has become popular across various domains, supporting a range of tasks and processes. However, due to the intrinsic complexity of LLMs, effective prompting is more…

The success of Large Language Models (LLMs) has led to a parallel rise in the development of Large Multimodal Models (LMMs), which have begun to transform a variety of applications. These sophisticated multimodal models are designed to…

Artificial Intelligence · Computer Science 2025-05-20 Fouad Trad , Ali Chehab

There is great interest in fine-tuning frontier large language models (LLMs) to inject new information and update existing knowledge. While commercial LLM fine-tuning APIs from providers such as OpenAI and Google promise flexible adaptation…

Computation and Language · Computer Science 2024-11-13 Eric Wu , Kevin Wu , James Zou

Traditional language models have been extensively evaluated for software engineering domain, however the potential of ChatGPT and Gemini have not been fully explored. To fulfill this gap, the paper in hand presents a comprehensive case…

Software Engineering · Computer Science 2024-12-03 Summra Saleem , Muhammad Nabeel Asim , Ludger Van Elst , Andreas Dengel

We present an outline of the first large language model (LLM) based chatbot application in the context of patient-reported outcome measures (PROMs) for diabetic retinopathy. By utilizing the capabilities of current LLMs, we enable patients…

Computation and Language · Computer Science 2024-11-06 Maren Pielka , Tobias Schneider , Jan Terheyden , Rafet Sifa

This study evaluates the capabilities of Multimodal Large Language Models (LLMs) and Vision Language Models (VLMs) in the task of single-label classification of Christian Iconography. The goal was to assess whether general-purpose VLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Gianmarco Spinaci , Lukas Klic , Giovanni Colavizza

Automatically interpreting CT scans can ease the workload of radiologists. However, this is challenging mainly due to the scarcity of adequate datasets and reference standards for evaluation. This study aims to bridge this gap by…

Artificial Intelligence · Computer Science 2024-06-19 Qingqing Zhu , Benjamin Hou , Tejas S. Mathai , Pritam Mukherjee , Qiao Jin , Xiuying Chen , Zhizheng Wang , Ruida Cheng , Ronald M. Summers , Zhiyong Lu

Background: Machine learning (ML) enhances gait analysis but often lacks the level of interpretability desired for clinical adoption. Large Language Models (LLMs) may offer explanatory capabilities and confidence-aware outputs when applied…

Machine Learning · Computer Science 2026-03-17 Carlo Dindorf , Jonas Dully , Rebecca Keilhauer , Michael Lorenz , Michael Fröhlich

Large Language Models (LLMs) are trained on massive amounts of data, enabling their application across diverse domains and tasks. Despite their remarkable performance, most LLMs are developed and evaluated primarily in English. Recently, a…

Computation and Language · Computer Science 2024-10-18 Krishno Dey , Prerona Tarannum , Md. Arid Hasan , Imran Razzak , Usman Naseem

This study explores the feasibility of using large language models (LLMs), specifically GPT-4o (ChatGPT), for automated grading of conceptual questions in an undergraduate Mechanical Engineering course. We compared the grading performance…

Computers and Society · Computer Science 2024-11-07 Rujun Gao , Xiaosu Guo , Xiaodi Li , Arun Balajiee Lekshmi Narayanan , Naveen Thomas , Arun R. Srinivasa

Large Language Models (LLMs) such as ChatGPT, Claude, and Gemini increasingly act as general-purpose copilots, yet they often respond with unnecessary length on simple requests, adding redundant explanations, hedging, or boilerplate that…

Machine Learning · Computer Science 2026-01-05 Vadim Borisov , Michael Gröger , Mina Mikhael , Richard H. Schreiber

We introduce a new approach in which several advanced large language models-specifically GPT-4-0125-preview, Meta-LLAMA-3-70B-Instruct, Claude-3-Opus, and Gemini-1.5-Flash-collaborate to both produce and answer intricate, doctoral-level…

We introduce AudioCapBench, a benchmark for evaluating audio captioning capabilities of large multimodal models. \method covers three distinct audio domains, including environmental sound, music, and speech, with 1,000 curated evaluation…

This paper presents an evaluation of three LLMs, GPT-4, Claude 3, and Gemini, for automated Behaviour-Driven Development (BDD) scenarios generation. To support this evaluation, we constructed a dataset of 500 user stories, requirement…

Software Engineering · Computer Science 2026-03-06 Amila Rathnayake , Mojtaba Shahin , Golnoush Abaei

This paper introduces ChatbotManip, a novel dataset for studying manipulation in Chatbots. It contains simulated generated conversations between a chatbot and a (simulated) user, where the chatbot is explicitly asked to showcase…

Computation and Language · Computer Science 2026-05-12 Jack Contro , Simrat Deol , Yulan He , Martim Brandão

Traffic safety remains a critical global concern, with timely and accurate accident detection essential for hazard reduction and rapid emergency response. Infrastructure-based vision sensors offer scalable and efficient solutions for…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Ilhan Skender , Kailin Tong , Selim Solmaz , Daniel Watzenig

Leveraging the power of multimodal large language models (LLMs) offers a promising approach to enhancing the accuracy and interpretability of morphing attack detection (MAD), especially in real-world biometric applications. This work…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Ria Shekhawat , Hailin Li , Raghavendra Ramachandra , Sushma Venkatesh

Multimodal Large Language Models (LLMs) claim "musical understanding" via evaluations that conflate listening with score reading. We benchmark three SOTA LLMs (Gemini 2.5 Pro, Gemini 2.5 Flash, and Qwen2.5-Omni) across three core music…

Sound · Computer Science 2025-10-28 Brandon James Carone , Iran R. Roman , Pablo Ripollés

This full research-to-practice paper explores approaches for developing course chatbots by comparing low-code platforms and custom-coded solutions in educational contexts. With the rise of Large Language Models (LLMs) like GPT-4 and LLaMA,…

Human-Computer Interaction · Computer Science 2025-09-18 Hemil Mehta , Tanvi Raut , Kohav Yadav , Edward F. Gehringer

This study presents a novel multi-model fusion framework leveraging two state-of-the-art large language models (LLMs), ChatGPT and Claude, to enhance the reliability of chest X-ray interpretation on the CheXpert dataset. From the full…

Computation and Language · Computer Science 2025-10-21 Md Kamrul Siam , Md Jobair Hossain Faruk , Jerry Q. Cheng , Huanying Gu
‹ Prev 1 8 9 10 Next ›