中文
相关论文

相关论文: CharacterEval: A Chinese Benchmark for Role-Playin…

200 篇论文

The advancement of Large Language Models (LLMs) has spurred significant interest in Role-Playing Agents (RPAs) for applications such as emotional companionship and virtual interaction. However, recent RPAs are often built on explicit…

计算与语言 · 计算机科学 2025-06-03 Yihong Tang , Kehai Chen , Muyun Yang , Zhengyu Niu , Jing Li , Tiejun Zhao , Min Zhang

State-of-the-art large language models (LLMs) are now claiming remarkable supported context lengths of 256k or even more. In contrast, the average context lengths of mainstream benchmarks are insufficient (5k-21k), and they suffer from…

计算与语言 · 计算机科学 2025-10-23 Tao Yuan , Xuefei Ning , Dong Zhou , Zhijie Yang , Shiyao Li , Minghui Zhuang , Zheyue Tan , Zhuyu Yao , Dahua Lin , Boxun Li , Guohao Dai , Shengen Yan , Yu Wang

Recent large language model (LLM)-driven chat assistant systems have integrated memory components to track user-assistant chat histories, enabling more accurate and personalized responses. However, their long-term memory capabilities in…

计算与语言 · 计算机科学 2025-03-06 Di Wu , Hongwei Wang , Wenhao Yu , Yuwei Zhang , Kai-Wei Chang , Dong Yu

In this paper, we present Edu-Values, the first Chinese education values evaluation benchmark that includes seven core values: professional philosophy, teachers' professional ethics, education laws and regulations, cultural literacy,…

计算与语言 · 计算机科学 2026-03-24 Peiyi Zhang , Yazhou Zhang , Bo Wang , Lu Rong , Prayag Tiwari , Jing Qin

Large language models (LLMs) hold great promise for assisting clinical interviews due to their fluent interactive capabilities and extensive medical knowledge. However, the lack of high-quality interview dialogue data and widely accepted…

计算与语言 · 计算机科学 2025-04-15 Jing Chen , Zhihua Wei , Wei Zhang , Yingying Hu , Qiong Zhang

In this work, we present the largest benchmark to date on linguistic acceptability: Multilingual Evaluation of Linguistic Acceptability -- MELA, with 46K samples covering 10 languages from a diverse set of language families. We establish…

计算与语言 · 计算机科学 2024-06-07 Ziyin Zhang , Yikang Liu , Weifang Huang , Junyu Mao , Rui Wang , Hai Hu

Recently, significant efforts have been devoted to enhancing the long-context capabilities of Large Language Models (LLMs), particularly in long-context reasoning. To facilitate this research, we propose \textbf{DetectiveQA}, a dataset…

计算与语言 · 计算机科学 2025-03-17 Zhe Xu , Jiasheng Ye , Xiaoran Liu , Xiangyang Liu , Tianxiang Sun , Zhigeng Liu , Qipeng Guo , Linlin Li , Qun Liu , Xuanjing Huang , Xipeng Qiu

The widespread adoption of large language models (LLMs) across various regions underscores the urgent need to evaluate their alignment with human values. Current benchmarks, however, fall short of effectively uncovering safety…

While LLMs enable personalized chatbots, their effectiveness in child-centered personalization remains unclear, as systematic evaluation of child-specific preferences is still lacking. To address this gap, we introduce ChildEval, a…

计算与语言 · 计算机科学 2026-05-28 Yanyan Luo , Xue Han , Chunxu Zhao , Ruiqiao Bai , Yaxing Zhang , Qian Hu , Lijun Mei , Junlan Feng

Different from the traditional translation tasks, classical Chinese poetry translation requires both adequacy and fluency in translating culturally and historically significant content and linguistic poetic elegance. Large language models…

计算与语言 · 计算机科学 2024-12-31 Andong Chen , Lianzhang Lou , Kehai Chen , Xuefeng Bai , Yang Xiang , Muyun Yang , Tiejun Zhao , Min Zhang

Large language models (LLMs) have shown impressive capabilities across various natural language tasks. However, evaluating their alignment with human preferences remains a challenge. To this end, we propose a comprehensive human evaluation…

We propose a new benchmark, ComperDial, which facilitates the training and evaluation of evaluation metrics for open-domain dialogue systems. ComperDial consists of human-scored responses for 10,395 dialogue turns in 1,485 conversations…

We propose OutboundEval, a comprehensive benchmark for evaluating large language models (LLMs) in expert-level intelligent outbound calling scenarios. Unlike existing methods that suffer from three key limitations - insufficient dataset…

To evaluate Large Language Models (LLMs) for question answering (QA), traditional methods typically focus on assessing single-turn responses to given questions. However, this approach doesn't capture the dynamic nature of human-AI…

计算与语言 · 计算机科学 2024-11-19 Ruosen Li , Ruochen Li , Barry Wang , Xinya Du

Legal Large Language Models (LLMs) have shown promise in providing legal consultations to non-experts. However, most existing Chinese legal consultation models are based on single-agent systems, which differ from real-world legal…

计算与语言 · 计算机科学 2024-12-17 Jingyun Sun , Chengxiao Dai , Zhongze Luo , Yangbo Chang , Yang Li

The recent introduction of the Assistants API highlights its potential for large language models (LLMs) in role-playing agents (RPA). However, maintaining consistent character personas remains a significant challenge due to variability in…

计算与语言 · 计算机科学 2025-02-25 Jeiyoon Park , Chanjun Park , Heuiseok Lim

The critical field of psychology necessitates a comprehensive benchmark to enhance the evaluation and development of domain-specific Large Language Models (LLMs). Existing MMLU-type benchmarks, such as C-EVAL and CMMLU, include…

计算与语言 · 计算机科学 2024-06-18 Junlei Zhang , Hongliang He , Nirui Song , Zhanchao Zhou , Shuyuan He , Shuai Zhang , Huachuan Qiu , Anqi Li , Yong Dai , Lizhi Ma , Zhenzhong Lan

Large language models (LLMs) demonstrate significant potential in advancing medical applications, yet their capabilities in addressing medical ethics challenges remain underexplored. This paper introduces MedEthicEval, a novel benchmark…

计算与语言 · 计算机科学 2025-03-05 Haoan Jin , Jiacheng Shi , Hanhui Xu , Kenny Q. Zhu , Mengyue Wu

Mental disorders are highly prevalent worldwide, but the shortage of psychiatrists and the inherent subjectivity of interview-based diagnosis create substantial barriers to timely and consistent mental-health assessment. Progress in…

多智能体系统 · 计算机科学 2026-02-12 Shihao Xu , Tiancheng Zhou , Jiatong Ma , Yanli Ding , Yiming Yan , Ming Xiao , Guoyi Li , Haiyang Geng , Yunyun Han , Jianhua Chen , Yafeng Deng

Recent advances in LLM-based role-playing language agents (RPLAs) have attracted broad attention in various applications. While chain-of-thought reasoning has shown importance in many tasks for LLMs, the internal thinking processes of RPLAs…

人工智能 · 计算机科学 2025-03-12 Rui Xu , MingYu Wang , XinTao Wang , Dakuan Lu , Xiaoyu Tan , Wei Chu , Yinghui Xu
‹ 上一页 1 8 9 10 下一页 ›