中文
相关论文

相关论文: Synthetic Users, Real Differences: an Evaluation F…

200 篇论文

Dialogue level quality estimation is vital for optimizing data driven dialogue management. Current automated methods to estimate turn and dialogue level user satisfaction employ hand-crafted features and rely on complex annotation schemes,…

计算与语言 · 计算机科学 2020-10-12 Praveen Kumar Bodigutla , Aditya Tiwari , Josep Valls Vargas , Lazaros Polymenakos , Spyros Matsoukas

Problem: Effective patient-centered communication is a core competency for physicians. However, both seasoned providers and medical trainees report decreased confidence in leading conversations on sensitive topics such as goals of care or…

人机交互 · 计算机科学 2024-05-31 Simon N. Chu , Alex J. Goodell

Speech-driven Talking Human (TH) generation, commonly known as "Talker," currently faces limitations in multi-subject driving capabilities. Extending this paradigm to "Multi-Talker," capable of animating multiple subjects simultaneously,…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Yingjie Zhou , Xilei Zhu , Siyu Ren , Ziyi Zhao , Ziwen Wang , Farong Wen , Yu Zhou , Jiezhang Cao , Xiongkuo Min , Fengjiao Chen , Xiaoyu Li , Xuezhi Cao , Guangtao Zhai , Xiaohong Liu

User interactions with LLMs are shaped by prior experiences and individual exploration, but in-lab studies do not provide system designers with visibility into these in-the-wild factors. This work explores a new approach to studying…

人机交互 · 计算机科学 2026-05-08 Shengqi Zhu , Jeffrey M. Rzeszotarski , David Mimno

Proposal of large-scale datasets has facilitated research on deep neural models for news summarization. Deep learning can also be potentially useful for spoken dialogue summarization, which can benefit a range of real-life scenarios…

计算与语言 · 计算机科学 2021-06-17 Yulong Chen , Yang Liu , Liang Chen , Yue Zhang

Existing function-calling benchmarks focus on single-turn interactions. However, they overlook the complexity of real-world scenarios. To quantify how existing benchmarks address practical applications, we introduce DICE-SCORE, a metric…

计算与语言 · 计算机科学 2025-07-03 Kyochul Jang , Donghyeon Lee , Kyusik Kim , Dongseok Heo , Taewhoo Lee , Woojeong Kim , Bongwon Suh

We introduce BotSIM, a modular, open-source Bot SIMulation environment with dialog generation, user simulation and conversation analytics capabilities. BotSIM aims to serve as a one-stop solution for large-scale data-efficient end-to-end…

计算与语言 · 计算机科学 2022-12-01 Guangsen Wang , Shafiq Joty , Junnan Li , Steven Hoi

Supervised fine-tuning with synthesized instructions has been a common practice for adapting LLMs to domain-specific QA tasks. However, the synthesized instructions deviate from real user questions and expected answers. This study proposes…

计算与语言 · 计算机科学 2025-02-14 Yang Li , Mingxuan Luo , Yeyun Gong , Chen Lin , Jian Jiao , Yi Liu , Kaili Huang

The aim of this workshop is two-fold. First, it aims to establish a research community focused on design and evaluation of synthetic speech (TTS) interfaces that are tailored not only to goal oriented tasks (e.g., food ordering, online…

人机交互 · 计算机科学 2024-02-27 Mateusz Dubiel , Matthew Aylett , Anuschka Schmitt , Zilin Ma , Gary Hsieh , Thiemo Wambsganss

A chatbot's personality design is key to interaction quality. As chatbots evolved from rule-based systems to those powered by large language models (LLMs), evaluating the effectiveness of their personality design has become increasingly…

人机交互 · 计算机科学 2025-08-11 Huiqi Zou , Pengda Wang , Zihan Yan , Tianjun Sun , Ziang Xiao

Social media has emerged as a cornerstone of social movements, wielding significant influence in driving societal change. Simulating the response of the public and forecasting the potential impact has become increasingly important. However,…

计算机与社会 · 计算机科学 2024-06-18 Xinyi Mou , Zhongyu Wei , Xuanjing Huang

A natural way to resolve different points of view and form opinions is through exchanging arguments and knowledge. Facing the vast amount of available information on the internet, people tend to focus on information consistent with their…

人机交互 · 计算机科学 2023-08-21 Klaus Weber , Annalena Aicher , Wolfang Minker , Stefan Ultes , Elisabeth André

This work discusses the benefits of having multiple simulated environments with different degrees of realism for the development of algorithms in scenarios populated by autonomous nodes capable of communication and mobility. This approach…

网络与互联网体系结构 · 计算机科学 2024-03-25 Thiago de Souza Lamenza , Josef Kamysek , Bruno Jose Olivieri de Souza , Markus Endler

Popular conversational agents frameworks such as Alexa Skills Kit (ASK) and Google Actions (gActions) offer unprecedented opportunities for facilitating the development and deployment of voice-enabled AI solutions in various verticals.…

计算与语言 · 计算机科学 2020-01-03 Walid Shalaby , Adriano Arantes , Teresa GonzalezDiaz , Chetan Gupta

With the advancement of large language models (LLMs), the focus in Conversational AI has shifted from merely generating coherent and relevant responses to tackling more complex challenges, such as personalizing dialogue systems. In an…

计算与语言 · 计算机科学 2025-02-13 Maria Molchanova , Anna Mikhailova , Anna Korzanova , Lidiia Ostyakova , Alexandra Dolidze

People increasingly hold sustained, open-ended conversations with large language models (LLMs). Public reports and early studies suggest that, in such settings, models can reinforce delusional or conspiratorial ideation or even amplify…

人机交互 · 计算机科学 2026-04-09 Peter Kirgis , Ben Hawriluk , Sherrie Feng , Aslan Bilimer , Sam Paech , Zeynep Tufekci

Dialogue summarization is abstractive in nature, making it suffer from factual errors. The factual correctness of summaries has the highest priority before practical applications. Many efforts have been made to improve faithfulness in text…

计算与语言 · 计算机科学 2022-10-24 Bin Wang , Chen Zhang , Yan Zhang , Yiming Chen , Haizhou Li

Progress in conversational information access (CIA) systems has been hindered by the difficulty of evaluating such systems with reproducible experiments. While user simulation offers a promising solution, the lack of infrastructure and…

信息检索 · 计算机科学 2025-10-27 Nolwenn Bernard , Sharath Chandra Etagi Suresh , Krisztian Balog , ChengXiang Zhai

Human evaluation is becoming a necessity to test the performance of Chatbots. However, off-the-shelf settings suffer the severe reliability and replication issues partly because of the extremely high diversity of criteria. It is high time…

计算与语言 · 计算机科学 2021-05-25 Hongru Liang , Huaqing Li

Despite growing interest in applications based on natural customer support conversations, there exist remarkably few publicly available datasets that reflect the expected characteristics of conversations in these settings. Existing…

计算与语言 · 计算机科学 2023-05-05 James Gung , Emily Moeng , Wesley Rose , Arshit Gupta , Yi Zhang , Saab Mansour