中文
相关论文

相关论文: EVA-Bench: A New End-to-end Framework for Evaluati…

200 篇论文

Automatic speech recognition (ASR) outcomes serve as input for downstream tasks, substantially impacting the satisfaction level of end-users. Hence, the diagnosis and enhancement of the vulnerabilities present in the ASR model bear…

计算与语言 · 计算机科学 2024-01-29 Seonmin Koo , Chanjun Park , Jinsung Kim , Jaehyung Seo , Sugyeong Eo , Hyeonseok Moon , Heuiseok Lim

Clinical AI systems require not just point-in-time evaluation but continuous governance: the ongoing practice of monitoring, evaluating, iterating, and re-evaluating performance throughout deployment. We present an end-to-end framework of…

Large Language Models (LLMs) based autonomous agents demonstrate multifaceted capabilities to contribute substantially to economic production. However, existing benchmarks remain focused on single agentic capability, failing to capture…

The rapid advancement of speech-to-speech (S2S) large language models (LLMs) has significantly improved real-time spoken interaction. However, current evaluation frameworks remain inadequate for assessing performance in complex, multi-turn…

计算与语言 · 计算机科学 2025-09-16 Yuhao Du , Qianwei Huang , Guo Zhu , Zhanchen Dai , Shunian Chen , Qiming Zhu , Le Pan , Minghao Chen , Yuhao Zhang , Li Zhou , Benyou Wang , Haizhou Li

Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models (Omni-LLMs), which are capable of processing multi-modal information including vision and audio,…

多媒体 · 计算机科学 2026-05-15 Jianghan Chao , Jianzhang Gao , Wenhui Tan , Yuchong Sun , Ruihua Song , Liyun Ru

Real-time voice assistants must revise task state when users interrupt mid-response, but existing spoken-dialog benchmarks largely evaluate turn-based interaction and miss this failure mode. We introduce EchoChain, a controlled benchmark…

计算与语言 · 计算机科学 2026-04-21 Smit Nautambhai Modi , Gandharv Mahajan , Marc Wetter , Randall Welles

Modern voice cloning, also known as zero-shot text-to-speech (TTS), can synthesize speech that closely matches a target speaker from only seconds of reference audio, enabling applications such as personalized speech interfaces and dubbing.…

声音 · 计算机科学 2026-05-26 Ruinan Jin , Xinting Liao , Hanlin Yu , Deval Pandya , Xiaoxiao Li

Recent end-to-end spoken dialogue models enable natural interaction. However, as user demands become increasingly complex, models that rely solely on conversational abilities often struggle to cope. Incorporating agentic capabilities is…

声音 · 计算机科学 2026-04-20 Tianle Liang , Yifu Chen , Shengpeng Ji , Yijun Chen , Zhiyang Jia , Jingyu Lu , Fan Zhuo , Xueyi Pu , Yangzhuo Li , Zhou Zhao

The human brain contextually exploits heterogeneous sensory information to efficiently perform cognitive tasks including vision and hearing. For example, during the cocktail party situation, the human auditory cortex contextually integrates…

声音 · 计算机科学 2021-12-17 Mandar Gogate , Kia Dashtipour , Amir Hussain

Audio-visual navigation of an agent towards locating an audio goal is a challenging task especially when the audio is sporadic or the environment is noisy. In this paper, we present CAVEN, a Conversation-based Audio-Visual Embodied…

计算机视觉与模式识别 · 计算机科学 2023-12-29 Xiulong Liu , Sudipta Paul , Moitreya Chatterjee , Anoop Cherian

Non-verbal vocalizations (NVVs) like laugh, sigh, and sob are essential for human-like speech, yet standardized evaluation remains limited in jointly assessing whether systems can generate the intended NVVs, place them correctly, and keep…

The rapid evolution of Large Language Models' has underscored the need for evaluation frameworks that are globally applicable, flexible, and modular, and that support a wide range of tasks, model types, and linguistic settings. We introduce…

计算与语言 · 计算机科学 2026-03-06 Samridhi Raj Sinha , Rajvee Sheth , Abhishek Upperwal , Mayank Singh

Speech large language models (SpeechLLMs) have extended human-machine interactions from the text modality to the dynamic speech domain. Spoken dialogues convey diverse information, including semantic concepts, acoustic variations,…

计算与语言 · 计算机科学 2026-01-14 Heyang Liu , Yuhao Wang , Ziyang Cheng , Hongcheng Liu , Yiqi Li , Yixuan Hou , Ronghua Wu , Qunshan Gu , Yanfeng Wang , Yu Wang

The emergence of "vibe coding" platforms, where users describe applications in natural language and AI agents autonomously generate full-stack software, has created a need for rigorous evaluation beyond code-level benchmarks. In order to…

多智能体系统 · 计算机科学 2026-05-07 Siddhant Saxena , Nilesh Trivedi , Vinayaka Jyothi

Large Language Models (LLMs) have shown strong potential as conversational agents. Yet, their effectiveness remains limited by deficiencies in robust long-term memory, particularly in complex, long-term web-based services such as online…

计算与语言 · 计算机科学 2026-02-03 Tiantian Chen , Jiaqi Lu , Ying Shen , Lin Zhang

Recent advances in audio-aware large language models have shown strong performance on audio question answering. However, existing benchmarks mainly cover answerable questions and overlook the challenge of unanswerable ones, where no…

音频与语音处理 · 电气工程与系统科学 2026-05-12 Chun-Yi Kuan , Hung-yi Lee

Recent significant advancements in Large Language Models (LLMs) have greatly propelled the development of Role-Playing Conversational Agents (RPCAs). These systems aim to create immersive user experiences through consistent persona…

计算与语言 · 计算机科学 2025-09-05 Weihao Wu , Liang Cao , Xinyu Wu , Zhiwei Lin , Rui Niu , Jingbei Li , Zhiyong Wu

Recently, instruction-following audio-language models have received broad attention for human-audio interaction. However, the absence of benchmarks capable of evaluating audio-centric interaction capabilities has impeded advancements in…

音频与语音处理 · 电气工程与系统科学 2024-07-29 Qian Yang , Jin Xu , Wenrui Liu , Yunfei Chu , Ziyue Jiang , Xiaohuan Zhou , Yichong Leng , Yuanjun Lv , Zhou Zhao , Chang Zhou , Jingren Zhou

The rapid advancement of Vision-Language Models (VLMs) has significantly advanced the development of Embodied Question Answering (EQA), enhancing agents' abilities in language understanding and reasoning within complex and realistic…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Tao Wu , Chuhao Zhou , Yen Heng Wong , Lin Gu , Jianfei Yang

We introduce AudioBench, a universal benchmark designed to evaluate Audio Large Language Models (AudioLLMs). It encompasses 8 distinct tasks and 26 datasets, among which, 7 are newly proposed datasets. The evaluation targets three main…

声音 · 计算机科学 2025-05-07 Bin Wang , Xunlong Zou , Geyu Lin , Shuo Sun , Zhuohan Liu , Wenyu Zhang , Zhengyuan Liu , AiTi Aw , Nancy F. Chen