English
Related papers

Related papers: TTA-Bench: A Comprehensive Benchmark for Evaluatin…

200 papers

Audio classifiers frequently face domain shift, when models trained on one dataset lose accuracy on data recorded in acoustically different conditions. Previous Test-Time Adaptation (TTA) research in speech and sound analysis often…

Sound · Computer Science 2025-11-25 Weichuang Shao , Iman Yi Liao , Tomas Henrique Bode Maul , Tissa Chandesa

We present MEGA-Bench, an evaluation suite that scales multimodal evaluation to over 500 real-world tasks, to address the highly heterogeneous daily use cases of end users. Our objective is to optimize for a set of high-quality data samples…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Jiacheng Chen , Tianhao Liang , Sherman Siu , Zhengqing Wang , Kai Wang , Yubo Wang , Yuansheng Ni , Wang Zhu , Ziyan Jiang , Bohan Lyu , Dongfu Jiang , Xuan He , Yuan Liu , Hexiang Hu , Xiang Yue , Wenhu Chen

Thanks to recent advancements in scalable deep architectures and large-scale pretraining, text-to-video generation has achieved unprecedented capabilities in producing high-fidelity, instruction-following content across a wide range of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Xuyang Guo , Jiayan Huo , Zhenmei Shi , Zhao Song , Jiahao Zhang , Jiale Zhao

Test-time adaptation (TTA) aims to boost the generalization capability of a trained model by conducting self-/unsupervised learning during the testing phase. While most existing TTA methods for video primarily utilize visual supervisory…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Runhao Zeng , Qi Deng , Ronghao Zhang , Shuaicheng Niu , Jian Chen , Xiping Hu , Victor C. M. Leung

We provide a new multi-task benchmark for evaluating text-to-image models. We perform a human evaluation comparing the most common open-source (Stable Diffusion) and commercial (DALL-E 2) models. Twenty computer science AI graduate students…

Modern voice cloning, also known as zero-shot text-to-speech (TTS), can synthesize speech that closely matches a target speaker from only seconds of reference audio, enabling applications such as personalized speech interfaces and dubbing.…

Sound · Computer Science 2026-05-26 Ruinan Jin , Xinting Liao , Hanlin Yu , Deval Pandya , Xiaoxiao Li

Text-to-image (T2I) models have garnered significant attention for generating high-quality images aligned with text prompts. However, rapid T2I model advancements reveal limitations in early benchmarks, lacking comprehensive evaluations,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Jingjing Chang , Yixiao Fang , Peng Xing , Shuhan Wu , Wei Cheng , Rui Wang , Xianfang Zeng , Gang Yu , Hai-Bao Chen

The development of audio foundation models has accelerated rapidly since the emergence of GPT-4o. However, the lack of comprehensive evaluation has become a critical bottleneck for further progress in the field, particularly in audio…

Think about how human handles complex reading tasks: marking key points, inferring their relationships, and structuring information to guide understanding and responses. Likewise, can a large language model benefit from text structure to…

Computation and Language · Computer Science 2026-03-05 Qinsi Wang , Hancheng Ye , Jinhee Kim , Jinghan Ke , Yifei Wang , Martin Kuo , Zishan Shao , Dongting Li , Yueqian Lin , Ting Jiang , Chiyue Wei , Qi Qian , Wei Wen , Helen Li , Yiran Chen

Character-based dialogue (aka role-playing) enables users to freely customize characters for interaction, which often relies on LLMs, raising the need to evaluate LLMs' character customization capability. However, existing benchmarks fail…

As large language models (LLMs) develop anthropomorphic abilities, they are increasingly being deployed as autonomous agents to interact with humans. However, evaluating their performance in realistic and complex social interactions remains…

Computation and Language · Computer Science 2025-10-28 Shuai Huang , Wenxuan Zhao , Jun Gao

Podcast script generation requires LLMs to synthesize structured, context-grounded dialogue from diverse inputs, yet systematic evaluation resources for this task remain limited. To bridge this gap, we introduce PodBench, a benchmark…

Computation and Language · Computer Science 2026-01-22 Chenning Xu , Mao Zheng , Mingyu Zheng , Mingyang Song

Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professional granularity to capture multi-dimensional aesthetic nuances. In this paper, we propose…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-30 Dapeng Wu , Shun Lei , Wei Tan , Guangzheng Li , Yunzhe Wang , Huaicheng Zhang , Lishi Zuo , Zhiyong Wu

Test-Time Adaptation (TTA) aims to adapt pre-trained models to the target domain during testing. In reality, this adaptability can be influenced by multiple factors. Researchers have identified various challenging scenarios and developed…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Chaoqun Du , Yulin Wang , Jiayi Guo , Yizeng Han , Jie Zhou , Gao Huang

Video generation has achieved remarkable progress, with generated videos increasingly resembling real ones. However, the rapid advance in generation has outpaced the development of adequate evaluation metrics. Currently, the assessment of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Nabyl Quignon , Baptiste Chopin , Yaohui Wang , Antitza Dantcheva

The rapid advancements of Text-to-Image (T2I) models have ushered in a new phase of AI-generated content, marked by their growing ability to interpret and follow user instructions. However, existing T2I model evaluation benchmarks fall…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Xinyu Wei , Jinrui Zhang , Zeqing Wang , Hongyang Wei , Zhen Guo , Lei Zhang

Video-to-audio generation (V2A) is of increasing importance in domains such as film post-production, AR/VR, and sound design, particularly for the creation of Foley sound effects synchronized with on-screen actions. Foley requires…

Sound · Computer Science 2025-11-25 Satvik Dixit , Koichi Saito , Zhi Zhong , Yuki Mitsufuji , Chris Donahue

Text-to-video (T2V) models have shown remarkable performance in generating visually reasonable scenes, while their capability to leverage world knowledge for ensuring semantic consistency and factual accuracy remains largely understudied.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Yubin Chen , Xuyang Guo , Zhenmei Shi , Zhao Song , Jiahao Zhang

Video generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Ziqi Huang , Fan Zhang , Xiaojie Xu , Yinan He , Jiashuo Yu , Ziyue Dong , Qianli Ma , Nattapol Chanpaisit , Chenyang Si , Yuming Jiang , Yaohui Wang , Xinyuan Chen , Ying-Cong Chen , Limin Wang , Dahua Lin , Yu Qiao , Ziwei Liu

Interactive world models are advancing rapidly, yet existing benchmarks cover only part of the required competencies, leaving no unified standard for systematic evaluation. To fill this gap, we introduce WBench, a comprehensive multi-turn…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Kaining Ying , Hengrui Hu , Siyu Ren , Jiamu Li , Fengjiao Chen , Ziwen Wang , Xuezhi Cao , Xunliang Cai , Henghui Ding