English
Related papers

Related papers: LEGOEval: An Open-Source Toolkit for Dialogue Syst…

200 papers

Large Language Models (LLMs) have revolutionized AI-generated content evaluation, with the LLM-as-a-Judge paradigm becoming increasingly popular. However, current single-LLM evaluation approaches face significant challenges, including…

Artificial Intelligence · Computer Science 2026-03-03 Yiyue Qian , Shinan Zhang , Yun Zhou , Haibo Ding , Diego Socolinsky , Yi Zhang

A dialogue is essentially a multi-turn interaction among interlocutors. Effective evaluation metrics should reflect the dynamics of such interaction. Existing automatic metrics are focused very much on the turn-level quality, while ignoring…

Computation and Language · Computer Science 2021-06-08 Chen Zhang , Yiming Chen , Luis Fernando D'Haro , Yan Zhang , Thomas Friedrichs , Grandee Lee , Haizhou Li

Recently, pre-trained large language models (LLMs) have shown impressive abilities in generating codes from natural language descriptions, repairing buggy codes, translating codes between languages, and retrieving relevant code segments.…

Computation and Language · Computer Science 2023-11-07 Mohammad Abdullah Matin Khan , M Saiful Bari , Xuan Long Do , Weishi Wang , Md Rizwan Parvez , Shafiq Joty

Large language models (LLMs) are advancing at an unprecedented pace globally, with regions increasingly adopting these models for applications in their primary language. Evaluation of these models in diverse linguistic environments,…

Benchmarking AI systems in multi-turn interactive scenarios is essential for understanding their practical capabilities in real-world applications. However, existing evaluation protocols are highly heterogeneous, differing significantly in…

Computation and Language · Computer Science 2026-03-25 Qi Jia , Haodong Zhao , Dun Pei , Xiujie Song , Shibo Wang , Zijian Chen , Zicheng Zhang , Xiangyang Zhu , Guangtao Zhai

We present FlagEvalMM, an open-source evaluation framework designed to comprehensively assess multimodal models across a diverse range of vision-language understanding and generation tasks, such as visual question answering,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Zheqi He , Yesheng Liu , Jing-shu Zheng , Xuejing Li , Jin-Ge Yao , Bowen Qin , Richeng Xuan , Xi Yang

We create a new task-oriented dialog platform (MEEP) where agents are given considerable freedom in terms of utterances and API calls, but are constrained to work within a push-button environment. We include facilities for collecting…

Computation and Language · Computer Science 2020-10-13 Arkady Arkhangorodsky , Amittai Axelrod , Christopher Chu , Scot Fang , Yiqi Huang , Ajay Nagesh , Xing Shi , Boliang Zhang , Kevin Knight

We present two comprehensive benchmarks to evaluate the performance of language models in coding assistance tasks, covering code writing, debugging, code review, and conceptual understanding. Our main contribution includes two curated…

Software Engineering · Computer Science 2024-12-10 Nidhish Shah , Zulkuf Genc , Dogu Araci

Interpretability tools that offer explanations in the form of a dialogue have demonstrated their efficacy in enhancing users' understanding (Slack et al., 2023; Shen et al., 2023), as one-off explanations may fall short in providing…

Computation and Language · Computer Science 2024-04-25 Qianli Wang , Tatiana Anikina , Nils Feldhus , Josef van Genabith , Leonhard Hennig , Sebastian Möller

Geospatial code generation is emerging as a key direction in the integration of artificial intelligence and geoscientific analysis. However, there remains a lack of standardized tools for automatic evaluation in this domain. To address this…

Software Engineering · Computer Science 2025-05-20 Shuyang Hou , Zhangxiao Shen , Huayi Wu , Jianyuan Liang , Haoyue Jiao , Yaxian Qing , Xiaopu Zhang , Xu Li , Zhipeng Gui , Xuefeng Guan , Longgang Xiang

Advancements in audio foundation models (FMs) have fueled interest in end-to-end (E2E) spoken dialogue systems, but different web interfaces for each system makes it challenging to compare and contrast them effectively. Motivated by this,…

Advancing social-scientific research of human-AI interaction dynamics and outcomes often requires researchers to deliver experiences with live large-language models (LLMs) to participants through online survey platforms. However, technical…

Human-Computer Interaction · Computer Science 2026-02-13 Jaime Banks , Jon Stromer-Galley , Samiksha Singh , Collin Capano

Systems interacting with humans, such as assistive robots or chatbots, are increasingly integrated into our society. To prevent these systems from causing social, legal, ethical, empathetic, or cultural (SLEEC) harms, normative requirements…

Computers and Society · Computer Science 2025-01-23 Kevin Kolyakov , Lina Marsso , Nick Feng , Junwei Quan , Marsha Chechik

LLM-based automatic survey systems are transforming how users acquire information from the web by integrating retrieval, organization, and content synthesis into end-to-end generation pipelines. While recent works focus on developing new…

Computation and Language · Computer Science 2025-12-03 Jiahao Zhao , Shuaixing Zhang , Nan Xu , Lei Wang

We present ConvLab, an open-source multi-domain end-to-end dialog system platform, that enables researchers to quickly set up experiments with reusable components and compare a large set of different approaches, ranging from conventional…

Computation and Language · Computer Science 2019-04-19 Sungjin Lee , Qi Zhu , Ryuichi Takanobu , Xiang Li , Yaoqin Zhang , Zheng Zhang , Jinchao Li , Baolin Peng , Xiujun Li , Minlie Huang , Jianfeng Gao

Recently, an increasing number of multimodal (text and audio) benchmarks have emerged, primarily focusing on evaluating models' understanding capability. However, exploration into assessing generative capabilities remains limited,…

Measuring empathy in conversation can be challenging, as empathy is a complex and multifaceted psychological construct that involves both cognitive and emotional components. Human evaluations can be subjective, leading to inconsistent…

Artificial Intelligence · Computer Science 2023-01-31 Bushra Amjad , Muhammad Zeeshan , Mirza Omer Beg

Recently, there has been increasing interest in using Large Language Models (LLMs) to construct complex multi-agent systems to perform tasks such as compiling literature reviews, drafting consumer reports, and planning vacations. Many tools…

Computation and Language · Computer Science 2024-11-06 Andrew Zhu , Liam Dugan , Chris Callison-Burch

To evaluate Large Language Models (LLMs) for question answering (QA), traditional methods typically focus on assessing single-turn responses to given questions. However, this approach doesn't capture the dynamic nature of human-AI…

Computation and Language · Computer Science 2024-11-19 Ruosen Li , Ruochen Li , Barry Wang , Xinya Du

Visual generation models have achieved remarkable progress in computer graphics applications but still face significant challenges in real-world deployment. Current assessment approaches for visual generation tasks typically follow an…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Xiaoyue Mi , Fan Tang , Juan Cao , Qiang Sheng , Ziyao Huang , Peng Li , Yang Liu , Tong-Yee Lee