We introduce LiveSecBench, a continuously updated safety benchmark specifically for Chinese-language LLM application scenarios. LiveSecBench constructs a high-quality and unique dataset through a pipeline that combines automated generation with human verification. By periodically releasing new versions to expand the dataset and update evaluation metrics, LiveSecBench provides a robust and up-to-date standard for AI safety. In this report, we introduce our second release v251215, which evaluates across five dimensions (Public Safety, Fairness & Bias, Privacy, Truthfulness, and Mental Health Safety.) We evaluate 57 representative LLMs using an ELO rating system, offering a leaderboard of the current state of Chinese LLM safety. The result is available at https://livesecbench.intokentech.cn/.
@article{arxiv.2511.02366,
title = {LiveSecBench: A Dynamic and Event-Driven Safety Benchmark for Chinese Language Model Applications},
author = {Yudong Li and Peiru Yang and Feng Huang and Zhongliang Yang and Kecheng Wang and Haitian Li and Baocheng Chen and Xingyu An and Ziyu Liu and Youdan Yang and Kejiang Chen and Sifang Wan and Xu Wang and Yufei Sun and Liyan Wu and Ruiqi Zhou and Wenya Wen and Xingchi Gu and Tianxin Zhang and Yue Gao and Yongfeng Huang},
journal= {arXiv preprint arXiv:2511.02366},
year = {2025}
}