中文
相关论文

相关论文: EduEVAL-DB: A Role-Based Dataset for Pedagogical R…

200 篇论文

Large language models (LLMs) have demonstrated remarkable advances in mathematical and logical reasoning, yet statistics, as a distinct and integrative discipline, remains underexplored in benchmarking efforts. To address this gap, we…

Automated release note generation addresses the challenge of documenting frequent software updates, where manual efforts are time-consuming and prone to human error. Although recent advances in language models further enhance this process,…

软件工程 · 计算机科学 2025-11-05 Qianru Meng , Zhaochun Ren , Joost Visser

As Large Language Models (LLMs) and generative AI become more widespread, the content safety risks associated with their use also increase. We find a notable deficiency in high-quality content safety datasets and benchmarks that…

机器学习 · 计算机科学 2024-09-12 Shaona Ghosh , Prasoon Varshney , Erick Galinkin , Christopher Parisien

One core capability of Large Language Models (LLMs) is to follow natural language instructions. However, the evaluation of such abilities is not standardized: Human evaluations are expensive, slow, and not objectively reproducible, while…

计算与语言 · 计算机科学 2023-11-15 Jeffrey Zhou , Tianjian Lu , Swaroop Mishra , Siddhartha Brahma , Sujoy Basu , Yi Luan , Denny Zhou , Le Hou

We propose and release a new vulnerable source code dataset. We curate the dataset by crawling security issue websites, extracting vulnerability-fixing commits and source codes from the corresponding projects. Our new dataset contains…

密码学与安全 · 计算机科学 2023-08-10 Yizheng Chen , Zhoujie Ding , Lamya Alowain , Xinyun Chen , David Wagner

Emotional Intelligence (EI) is a critical yet underexplored dimension in the development of human-aligned LLMs. To address this gap, we introduce a unified, psychologically grounded four-layer taxonomy of EI tailored for large language…

计算与语言 · 计算机科学 2025-08-11 Nizi Nazar , Ehsaneddin Asgari

Deceptive behavior in AI systems is no longer theoretical: large language models strategically mislead without producing false statements, maintain deceptive strategies through safety training, and coordinate deception in multi-agent…

计算机与社会 · 计算机科学 2026-04-07 Jason Starace , Bert Baumgaertner , Terence Soule

Existing Grammatical Error Correction (GEC) systems suffer from limited reference diversity, leading to underestimated evaluation and restricted model generalization. To address this issue, we introduce the Judge of Edit-Level Validity…

计算与语言 · 计算机科学 2025-12-09 Yuhao Zhan , Yuqing Zhang , Jing Yuan , Qixiang Ma , Zhiqi Yang , Yu Gu , Zemin Liu , Fei Wu

As language models master existing reasoning benchmarks, we need new challenges to evaluate their cognitive frontiers. Puzzle-solving events are rich repositories of challenging multimodal problems that test a wide range of advanced…

Verbal confidence elicitation is widely used to extract uncertainty estimates from LLMs. We tested whether seven instruction-tuned open-weight models (3-9B parameters, four families) produce verbalised confidence that meets minimal validity…

计算与语言 · 计算机科学 2026-04-27 Jon-Paul Cacioli

While LLMs enable personalized chatbots, their effectiveness in child-centered personalization remains unclear, as systematic evaluation of child-specific preferences is still lacking. To address this gap, we introduce ChildEval, a…

计算与语言 · 计算机科学 2026-05-28 Yanyan Luo , Xue Han , Chunxu Zhao , Ruiqiao Bai , Yaxing Zhang , Qian Hu , Lijun Mei , Junlan Feng

Data contamination poses a significant challenge to the fairness of LLM evaluations in natural language processing tasks by inadvertently exposing models to test data during training. Current studies attempt to mitigate this issue by…

计算与语言 · 计算机科学 2025-11-25 Jingqian Zhao , Bingbing Wang , Geng Tu , Yice Zhang , Qianlong Wang , Bin Liang , Jing Li , Ruifeng Xu

In recent years, trustworthiness has garnered increasing attention and exploration in the field of intelligent education, due to the inherent sensitivity of educational scenarios, such as involving minors and vulnerable groups, highly…

计算机与社会 · 计算机科学 2026-01-30 Xiaoshan Yu , Shangshang Yang , Ziwen Wang , Haiping Ma , Xingyi Zhang

Recent advancements in Large Language Models (LLMs) and their increased accessibility have made it easier than ever for students to automatically generate texts, posing new challenges for educational institutions. To enforce norms of…

计算与语言 · 计算机科学 2025-08-12 Lukas Gehring , Benjamin Paaßen

EduChat (https://www.educhat.top/) is a large-scale language model (LLM)-based chatbot system in the education domain. Its goal is to support personalized, fair, and compassionate intelligent education, serving teachers, students, and…

Obtaining accurate class labels is often costly or unreliable, and may also be limited by privacy or other practical conditions. Compared with asking an annotator to provide the exact class, it is often easier to ask whether the true label…

机器学习 · 计算机科学 2026-05-11 Jiaxu Su , Junpeng Li , Changchun Hua , Yana Yang

Curated datasets for healthcare are often limited due to the need of human annotations from experts. In this paper, we present MedEval, a multi-level, multi-task, and multi-domain medical benchmark to facilitate the development of language…

计算与语言 · 计算机科学 2023-11-16 Zexue He , Yu Wang , An Yan , Yao Liu , Eric Y. Chang , Amilcare Gentili , Julian McAuley , Chun-Nan Hsu

The impact of software vulnerabilities on everyday software systems is significant. Despite deep learning models being proposed for vulnerability detection, their reliability is questionable. Prior evaluations show high recall/F1 scores of…

Analyzing the reasoning patterns of students in science classrooms is critical for understanding knowledge construction mechanism and improving instructional practice to maximize cognitive engagement, yet manual coding of classroom…

计算与语言 · 计算机科学 2026-05-08 Jiho Noh , Mukhesh Raghava Katragadda , Raymond Carl , Soon Lee

Text-driven video editing is rapidly advancing, yet its rigorous evaluation remains challenging due to the absence of dedicated video quality assessment (VQA) models capable of discerning the nuances of editing quality. To address this…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Juntong Wang , Jiarui Wang , Huiyu Duan , Guangtao Zhai , Xiongkuo Min