English
Related papers

Related papers: MPCEval: A Benchmark for Multi-Party Conversation …

200 papers

Audio-driven human animation methods, such as talking head and talking body generation, have made remarkable progress in generating synchronized facial movements and appealing visual quality videos. However, existing methods primarily focus…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Zhe Kong , Feng Gao , Yong Zhang , Zhuoliang Kang , Xiaoming Wei , Xunliang Cai , Guanying Chen , Wenhan Luo

We present new benchmarks on evaluation code generation models: MBXP and Multilingual HumanEval, and MathQA-X. These datasets cover over 10 programming languages and are generated using a scalable conversion framework that transpiles…

With the recent advancements in reasoning capabilities, tool calling using MCP servers and Audio Language Models (ALMs), development and integration of multi-modal agents (with voice and text support) has come to the industry forefront.…

Sound · Computer Science 2026-03-03 Anupam Purwar , Aditya Choudhary

To develop a reliable AI for psychological assessment, we introduce \texttt{PsychEval}, a multi-session, multi-therapy, and highly realistic benchmark designed to address three key challenges: \textbf{1) Can we train a highly realistic AI…

Artificial Intelligence · Computer Science 2026-01-09 Qianjun Pan , Junyi Wang , Jie Zhou , Yutao Yang , Junsong Li , Kaiyin Xu , Yougen Zhou , Yihan Li , Jingyuan Zhao , Qin Chen , Ningning Zhou , Kai Chen , Liang He

Written Multi-Party Conversations (WMPCs) are widely studied across disciplines, with social media as a primary data source due to their accessibility. However, these datasets raise privacy concerns and often reflect platform-specific…

Computation and Language · Computer Science 2026-03-30 Nicolò Penzo , Marco Guerini , Bruno Lepri , Goran Glavaš , Sara Tonelli

Recent advances in text-to-audio-video (T2AV) generation have enabled models to synthesize audio-visual videos with multi-participant dialogues. However, existing evaluation benchmarks remain largely designed for human-recorded videos or…

Handling multi-party dialogues represents a significant step for advancing spoken dialogue systems, necessitating the development of tasks specific to multi-party interactions. To address this challenge, we are constructing a multi-modal…

Computation and Language · Computer Science 2025-03-19 Koji Inoue , Divesh Lala , Mikey Elmers , Keiko Ochi , Tatsuya Kawahara

Video generation is rapidly evolving from single-shot synthesis to complex multi-shot audio-video (MSAV) narratives to meet real-world demands. However, evaluating such frontier models remains a fundamental challenge. Existing benchmarks…

Large Language Models (LLMs) demonstrate a notable capacity for adopting personas and engaging in role-playing. However, evaluating this ability presents significant challenges, as human assessments are resource-intensive and automated…

Computation and Language · Computer Science 2025-05-20 Yassine El Boudouri , Walter Nuninger , Julian Alvarez , Yvan Peter

The rapid advancements in large language models (LLMs) have presented challenges in evaluating those models. Existing evaluation methods are either reference-based or preference based, which inevitably need human intervention or introduce…

Computation and Language · Computer Science 2023-08-22 Dan Qiao , Chenfei Wu , Yaobo Liang , Juntao Li , Nan Duan

Multi-party dialogue generation presents significant challenges due to the complex interplay of multiple speakers and interwoven conversational threads. Traditional approaches often fall short in capturing these complexities, particularly…

Computation and Language · Computer Science 2025-03-13 Tianyu Sun , Kun Qian , Wenhong Wang

The growing capabilities of large language models and multimodal systems have spurred interest in voice-first AI assistants, yet existing benchmarks are inadequate for evaluating the full range of these systems' capabilities. We introduce…

Computation and Language · Computer Science 2025-09-29 Ke Wang , Houxing Ren , Zimu Lu , Mingjie Zhan , Hongsheng Li

Realistic practice and tailored feedback are key processes for training peer counselors with clinical skills. However, existing mechanisms of providing feedback largely rely on human supervision. Peer counselors often lack mechanisms to…

Computation and Language · Computer Science 2024-03-26 Alicja Chaszczewicz , Raj Sanjay Shah , Ryan Louie , Bruce A Arnow , Robert Kraut , Diyi Yang

Today, conversational systems are expected to handle conversations in multi-party settings, especially within Socially Assistive Robots (SARs). However, practical usability remains difficult as there are additional challenges to overcome,…

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from…

Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align reliably with human evaluation, especially for complex tasks that involve multiple modalities. To address this, we…

Artificial Intelligence · Computer Science 2025-05-26 Jihan Yao , Yushi Hu , Yujie Yi , Bin Han , Shangbin Feng , Guang Yang , Bingbing Wen , Ranjay Krishna , Lucy Lu Wang , Yulia Tsvetkov , Noah A. Smith , Banghua Zhu

Parliamentary speech generation presents specific challenges for large language models beyond standard text generation tasks. Unlike general text generation, parliamentary speeches require not only linguistic quality but also political…

Computation and Language · Computer Science 2025-11-12 Marios Koniaris , Argyro Tsipi , Panayiotis Tsanakas

The quality of daily spontaneous conversations is of importance towards both our well-being as well as the development of interactive social agents. Prior research directly studying the quality of social conversations has operationalized it…

Human-Computer Interaction · Computer Science 2022-07-14 Chirag Raman , Navin Raj Prabhu , Hayley Hung

State of the art large language models rely on randomization to respond to a prompt. As an immediate consequence, a model may respond differently to the same prompt if asked multiple times. In this work, we argue that the evaluation and…

Large Language Models (LLMs) have revolutionized AI-generated content evaluation, with the LLM-as-a-Judge paradigm becoming increasingly popular. However, current single-LLM evaluation approaches face significant challenges, including…

Artificial Intelligence · Computer Science 2026-03-03 Yiyue Qian , Shinan Zhang , Yun Zhou , Haibo Ding , Diego Socolinsky , Yi Zhang