中文
相关论文

相关论文: Synthetic Users, Real Differences: an Evaluation F…

200 篇论文

Tool-Augmented Language Models (TALMs) leverage external APIs to answer user queries across various domains. However, existing benchmark datasets for TALM research often feature simplistic dialogues that do not reflect real-world scenarios,…

计算与语言 · 计算机科学 2025-03-04 Jeonghoon Shim , Gyuhyeon Seo , Cheongsu Lim , Yohan Jo

This study investigates the development and assessment of an artificial human designed as a conversational AI chatbot, focusing on its role as a clinical psychologist. The project involved creating a specialized chatbot using the…

人机交互 · 计算机科学 2025-03-24 Birger Moell

Large Language Models (LLMs) are increasingly used to simulate human users in interactive settings such as therapy, education, and social role-play. While these simulations enable scalable training and evaluation of AI agents, off-the-shelf…

计算与语言 · 计算机科学 2025-11-04 Marwa Abdulhai , Ryan Cheng , Donovan Clay , Tim Althoff , Sergey Levine , Natasha Jaques

Realistic user simulation is crucial for training and evaluating multi-turn dialogue systems, yet creating simulators that accurately replicate human behavior remains a significant challenge. An effective simulator must expose the failure…

Next generation task-oriented dialog systems need to understand conversational contexts with their perceived surroundings, to effectively help users in the real-world multimodal environment. Existing task-oriented dialog datasets aimed…

计算与语言 · 计算机科学 2021-10-22 Satwik Kottur , Seungwhan Moon , Alborz Geramifard , Babak Damavandi

With ChatGPT-like large language models (LLM) prevailing in the community, how to evaluate the ability of LLMs is an open question. Existing evaluation methods suffer from following shortcomings: (1) constrained evaluation abilities, (2)…

人工智能 · 计算机科学 2023-08-09 Jiaju Lin , Haoran Zhao , Aochi Zhang , Yiting Wu , Huqiuyue Ping , Qin Chen

One of the difficulties in training dialogue systems is the lack of training data. We explore the possibility of creating dialogue data through the interaction between a dialogue system and a user simulator. Our goal is to develop a…

计算与语言 · 计算机科学 2021-07-27 Bo-Hsiang Tseng , Yinpei Dai , Florian Kreyssig , Bill Byrne

Large language models (LLMs) are increasingly used for mental-health support; yet prevailing evaluation methods--fluency metrics, preference tests, and generic dialogue benchmarks--fail to capture the clinically critical dimensions of…

Simulation is pivotal in evaluating the performance of autonomous driving systems due to the advantages of high efficiency and low cost compared to on-road testing. Bridging the gap between simulation and the real world requires realistic…

机器人学 · 计算机科学 2025-06-17 Haojie Xin , Xiaodong Zhang , Renzhi Tang , Songyang Yan , Qianrui Zhao , Chunze Yang , Wen Cui , Zijiang Yang

The unparalleled performance of closed-sourced ChatGPT has sparked efforts towards its democratization, with notable strides made by leveraging real user and ChatGPT dialogues, as evidenced by Vicuna. However, due to challenges in gathering…

计算与语言 · 计算机科学 2024-08-27 Chuyi Kong , Yaxin Fan , Xiang Wan , Feng Jiang , Benyou Wang

Recent advancements have expanded the role of Large Language Models in board games from playing agents to creative co-designers. However, a critical gap remains: current systems lack the capacity to offer constructive critique grounded in…

人机交互 · 计算机科学 2026-04-14 Zizhen Li , Chuanhao Li , Yibin Wang , Yukang Feng , Jianwen Sun , Jiaxin Ai , Fanrui Zhang , Mingzhu Sun , Yifei Huang , Kaipeng Zhang

$ $Dialogue systems are evaluated depending on their type and purpose. Two categories are often distinguished: (1) task-oriented dialogue systems (TDS), which are typically evaluated on utility, i.e., their ability to complete a specified…

信息检索 · 计算机科学 2022-04-27 Clemencia Siro , Mohammad Aliannejadi , Maarten de Rijke

The sense of realism in avatar animation is a widely pursued goal in social VR applications. A common approach to enhancing realism is improving the match between avatar motion and real-world human movement. However, experience with…

人机交互 · 计算机科学 2025-09-22 Yudong Huang , Avneet Singh , Mark Roman Miller

Objective: This study aims to develop and validate an evaluation framework to ensure the safety and reliability of mental health chatbots, which are increasingly popular due to their accessibility, human-like interactions, and context-aware…

This study aims to understand users' perceptions of using the Dialogflow framework and verify the relationships among service awareness, task-technology fit, output quality, and TAM variables. Generalized Structured Component Analysis was…

人机交互 · 计算机科学 2023-02-09 Vinh T. Nguyen , Chuyen T. H. Nguyen

Dialog summarization has become increasingly important in managing and comprehending large-scale conversations across various domains. This task presents unique challenges in capturing the key points, context, and nuances of multi-turn long…

计算与语言 · 计算机科学 2024-02-28 Ankan Mullick , Ayan Kumar Bhowmick , Raghav R , Ravi Kokku , Prasenjit Dey , Pawan Goyal , Niloy Ganguly

An automated metric to evaluate dialogue quality is vital for optimizing data driven dialogue management. The common approach of relying on explicit user feedback during a conversation is intrusive and sparse. Current models to estimate…

机器学习 · 计算机科学 2019-11-21 Praveen Kumar Bodigutla , Lazaros Polymenakos , Spyros Matsoukas

Despite significant advancements in conversational AI, large language model (LLM)-powered chatbots often struggle with personalizing their responses according to individual user characteristics, such as technical expertise, learning style,…

人工智能 · 计算机科学 2025-06-18 Shahaf David , Yair Meidan , Ido Hersko , Daniel Varnovitzky , Dudu Mimran , Yuval Elovici , Asaf Shabtai

Due to the difficulty of acquiring extensive real-world data, robot simulation has become crucial for parallel training and sim-to-real transfer, highlighting the importance of scalable simulated robotic tasks. Foundation models have…

机器人学 · 计算机科学 2024-10-11 Feng Chen , Botian Xu , Pu Hua , Peiqi Duan , Yanchao Yang , Yi Ma , Huazhe Xu

We report the results of DialogSum Challenge, the shared task on summarizing real-life scenario dialogues at INLG 2022. Four teams participate in this shared task and three submit their system reports, exploring different methods to improve…

计算与语言 · 计算机科学 2022-09-07 Yulong Chen , Naihao Deng , Yang Liu , Yue Zhang