English

X-TURING: Towards an Enhanced and Efficient Turing Test for Long-Term Dialogue Agents

Computation and Language 2025-05-30 v2 Artificial Intelligence

Abstract

The Turing test examines whether AIs exhibit human-like behaviour in natural language conversations. The traditional setting limits each participant to one message at a time and requires constant human participation. This fails to reflect a natural conversational style and hinders the evaluation of dialogue agents based on Large Language Models (LLMs) in complex and prolonged interactions. This paper proposes \textbf{\textsc{X-Turing}}, which enhances the original test with a \textit{burst dialogue} pattern, allowing more dynamic exchanges using consecutive messages. It further reduces human workload by iteratively generating dialogues that simulate the long-term interaction between the agent and a human to compose the majority of the test process. With the \textit{pseudo-dialogue} history, the agent then engages in a shorter dialogue with a real human, which is paired with a human-human conversation on the same topic to be judged using questionnaires. We introduce the \textit{X-Turn Pass-Rate} metric to assess the human likeness of LLMs across varying durations. While LLMs like GPT-4 initially perform well, achieving pass rates of 51.9\% and 38.9\% during 3 turns and 10 turns of dialogues respectively, their performance drops as the dialogue progresses, which underscores the difficulty in maintaining consistency in the long term.

Keywords

Cite

@article{arxiv.2408.09853,
  title  = {X-TURING: Towards an Enhanced and Efficient Turing Test for Long-Term Dialogue Agents},
  author = {Weiqi Wu and Hongqiu Wu and Hai Zhao},
  journal= {arXiv preprint arXiv:2408.09853},
  year   = {2025}
}

Comments

Accepted to ACL 2025 Main Conference

R2 v1 2026-06-28T18:16:32.699Z