中文
相关论文

相关论文: TTA-Bench: A Comprehensive Benchmark for Evaluatin…

200 篇论文

With the rapid integration of advanced reasoning capabilities into spoken dialogue models, the field urgently demands benchmarks that transcend simple interactions to address real-world complexity. However, current evaluations predominantly…

计算与语言 · 计算机科学 2026-02-16 Yangzhuo Li , Shengpeng Ji , Yifu Chen , Tianle Liang , Haorong Ying , Yule Wang , Junbo Li , Jun Fang , Zhou Zhao

Recent advances in text-to-audio-video (T2AV) generation have enabled models to synthesize audio-visual videos with multi-participant dialogues. However, existing evaluation benchmarks remain largely designed for human-recorded videos or…

We present HealthBench, an open-source benchmark measuring the performance and safety of large language models in healthcare. HealthBench consists of 5,000 multi-turn conversations between a model and an individual user or healthcare…

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also neglect this…

计算机视觉与模式识别 · 计算机科学 2025-01-16 Kaiyue Sun , Kaiyi Huang , Xian Liu , Yue Wu , Zihan Xu , Zhenguo Li , Xihui Liu

Text-to-3D (T23D) generation has emerged as a crucial visual generation task, aiming at synthesizing 3D content from textual descriptions. Studies of this task are currently shifting from per-scene T23D, which requires optimization of the…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Xiao Cai , Sitong Su , Jingkuan Song , Pengpeng Zeng , Ji Zhang , Qinhong Du , Mengqi Li , Heng Tao Shen , Lianli Gao

Recent advances in Text-to-3D (T23D) generative models have enabled the synthesis of diverse, high-fidelity 3D assets from textual prompts. However, existing challenges restrict the development of reliable T23D quality assessment (T23DQA).…

计算机视觉与模式识别 · 计算机科学 2025-11-06 Bingyang Cui , Yujie Zhang , Qi Yang , Zhu Li , Yiling Xu

We introduce AudioBench, a universal benchmark designed to evaluate Audio Large Language Models (AudioLLMs). It encompasses 8 distinct tasks and 26 datasets, among which, 7 are newly proposed datasets. The evaluation targets three main…

声音 · 计算机科学 2025-05-07 Bin Wang , Xunlong Zou , Geyu Lin , Shuo Sun , Zhuohan Liu , Wenyu Zhang , Zhengyuan Liu , AiTi Aw , Nancy F. Chen

Text-to-video (T2V) generation models have made significant progress in creating visually appealing videos. However, they struggle with generating coherent sequential narratives that require logical progression through multiple events.…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Zhengxu Tang , Zizheng Wang , Luning Wang , Zitao Shuai , Chenhao Zhang , Siyu Qian , Yirui Wu , Bohao Wang , Haosong Rao , Zhenyu Yang , Chenwei Wu

As conversational AI-based dialogue management has increasingly become a trending topic, the need for a standardized and reliable evaluation procedure grows even more pressing. The current state of affairs suggests various evaluation…

计算与语言 · 计算机科学 2020-06-12 Sarah E. Finch , Jinho D. Choi

As large language models continue to advance, their application in educational contexts remains underexplored and under-optimized. In this paper, we address this gap by introducing the first diverse benchmark tailored for educational…

Generative models have driven significant progress in a variety of AI tasks, including text-to-video generation, where models like Video LDM and Stable Video Diffusion can produce realistic, movie-level videos from textual instructions.…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Xuyang Guo , Zekai Huang , Jiayan Huo , Yingyu Liang , Zhenmei Shi , Zhao Song , Jiahao Zhang

Continual post-training adapts a single text-to-image diffusion model to learn new tasks without incurring the cost of separate models, but naive post-training causes forgetting of pretrained knowledge and undermines zero-shot…

计算机视觉与模式识别 · 计算机科学 2025-05-23 Zhehao Huang , Yuhang Liu , Yixin Lou , Zhengbao He , Mingzhen He , Wenxing Zhou , Tao Li , Kehan Li , Zeyi Huang , Xiaolin Huang

The rapid advancement of AIGC-based video generation has underscored the critical need for comprehensive evaluation frameworks that go beyond traditional generation quality metrics to encompass aesthetic appeal. However, existing benchmarks…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Longteng Jiang , DanDan Zheng , Qianqian Qiao , Heng Huang , Huaye Wang , Yihang Bo , Bao Peng , Jingdong Chen , Jun Zhou , Xin Jin

Movie dubbing has advanced significantly, yet assessing the real-world effectiveness of these models remains challenging. A comprehensive evaluation benchmark is crucial for two key reasons: 1) Existing metrics fail to fully capture the…

机器学习 · 计算机科学 2025-05-06 Chaoyi Wang , Junjie Zheng , Zihao Chen , Shiyu Xia , Chaofan Ding , Xiaohao Zhang , Xi Tao , Xiaoming He , Xinhan Di

We propose a novel objective evaluation metric for synthesized audio in text-to-audio (TTA), aiming to improve the performance of TTA models. In TTA, subjective evaluation of the synthesized sound is an important, but its implementation…

声音 · 计算机科学 2025-07-02 Minoru Kishi , Ryosuke Sakai , Shinnosuke Takamichi , Yusuke Kanamori , Yuki Okamoto

Despite the impressive advances in text-to-image models, they often struggle to effectively compose complex scenes with multiple objects, displaying various attributes and relationships. To address this challenge, we present…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Kaiyi Huang , Chengqi Duan , Kaiyue Sun , Enze Xie , Zhenguo Li , Xihui Liu

Recently, instruction-following audio-language models have received broad attention for human-audio interaction. However, the absence of benchmarks capable of evaluating audio-centric interaction capabilities has impeded advancements in…

音频与语音处理 · 电气工程与系统科学 2024-07-29 Qian Yang , Jin Xu , Wenrui Liu , Yunfei Chu , Ziyue Jiang , Xiaohuan Zhou , Yichong Leng , Yuanjun Lv , Zhou Zhao , Chang Zhou , Jingren Zhou

Text-to-audio (TTA) generation can significantly benefit the media industry by reducing production costs and enhancing work efficiency. However, most current TTA models (primarily diffusion-based) suffer from slow inference speeds and high…

声音 · 计算机科学 2025-12-30 HaeChun Chung

Text-to-Audio (TTA) aims to generate audio that corresponds to the given text description, playing a crucial role in media production. The text descriptions in TTA datasets lack rich variations and diversity, resulting in a drop in TTA…

Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align with human perceptions, and effective quantitative metrics…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Shangkun Sun , Xiaoyu Liang , Songlin Fan , Wenxu Gao , Wei Gao