中文
相关论文

相关论文: EduEVAL-DB: A Role-Based Dataset for Pedagogical R…

200 篇论文

The rapid proliferation of benchmarks has created significant challenges in reproducibility, transparency, and informed decision-making. However, unlike datasets and models -- which benefit from structured documentation frameworks like…

机器学习 · 计算机科学 2025-12-04 Florian Bordes , Candace Ross , Justine T Kao , Evangelia Spiliopoulou , Adina Williams

Text-to-audio (TTA) generation is advancing rapidly, but evaluation remains challenging because human listening studies are expensive and existing automatic metrics capture only limited aspects of perceptual quality. We introduce AudioEval,…

声音 · 计算机科学 2026-01-30 Hui Wang , Jinghua Zhao , Junyang Cheng , Cheng Liu , Yuhang Jia , Haoqin Sun , Jiaming Zhou , Yong Qin

In the age of artificial intelligence (AI), providing learners with suitable and sufficient explanations of AI-based recommendation algorithm's output becomes essential to enable them to make an informed decision about it. However, the…

人机交互 · 计算机科学 2024-02-14 Hasan Abu-Rasheed , Christian Weber , Madjid Fathi

DeepResearch agents represent a transformative AI paradigm, conducting expert-level research through sophisticated reasoning and multi-tool integration. However, evaluating these systems remains critically challenging due to open-ended…

人工智能 · 计算机科学 2025-10-10 Tianyu Fan , Xinyao Niu , Yuxiang Zheng , Fengji Zhang , Chengen Huang , Bei Chen , Junyang Lin , Chao Huang

The performance and usability of Large-Language Models (LLMs) are driving their use in explanation generation tasks. However, despite their widespread adoption, LLM explanations have been found to be unreliable, making it difficult for…

Rubric-based evaluation has become a prevailing paradigm for evaluating instruction following in large language models (LLMs). Despite its widespread use, the reliability of these rubric-level evaluations remains unclear, calling for…

人工智能 · 计算机科学 2026-03-27 Tianjun Pan , Xuan Lin , Wenyan Yang , Qianyu He , Shisong Chen , Licai Qi , Wanqing Xu , Hongwei Feng , Bo Xu , Yanghua Xiao

We present MSA-MathEval, our submission to the BEA 2025 Shared Task on evaluating AI tutor responses across four instructional dimensions: Mistake Identification, Mistake Location, Providing Guidance, and Actionability. Our approach uses a…

计算与语言 · 计算机科学 2025-05-27 Baraa Hikal , Mohamed Basem , Islam Oshallah , Ali Hamdi

Video generation models (VGMs) are rapidly entering classrooms, yet existing benchmarks evaluate only perceptual quality, intrinsic faithfulness, generic safety, or video as a reasoning medium, and none assesses whether the outputs are…

Recently, there has been increasing interest in transparency and interpretability in Deep Reinforcement Learning (DRL) systems. Verbal explanations, as the most natural way of communication in our daily life, deserve more attention, since…

人工智能 · 计算机科学 2020-12-25 Xinzhi Wang , Huao Li , Hui Zhang , Michael Lewis , Katia Sycara

Emotional support is a core capability in human-AI interaction, with applications including psychological counseling, role play, and companionship. However, existing evaluations of large language models (LLMs) often rely on short, static…

Large language models (LLMs) have a transformative impact on a variety of scientific tasks across disciplines including biology, chemistry, medicine, and physics. However, ensuring the safety alignment of these models in scientific research…

As Vision-Language Models (VLMs) become integral to educational decision-making, ensuring their fairness is paramount. However, current text-centric evaluations neglect the visual modality, leaving an unregulated channel for latent social…

人工智能 · 计算机科学 2026-04-15 Ruijia Li , Mingzi Zhang , Zengyi Yu , Yuang Wei , Bo Jiang

This concluding chapter explores how artificial intelligence (AI) is reshaping the purposes, practices, and outcomes of science education, and proposes a human-centered framework for its responsible integration. Drawing on insights from…

计算机与社会 · 计算机科学 2026-02-24 Xiaoming Zhai , Kent Crippen

Relational databases (RDBs) are widely regarded as the gold standard for storing structured information. Consequently, predictive tasks leveraging this data format hold significant application promise. Recently, Relational Deep Learning…

机器学习 · 计算机科学 2025-12-15 Jakub Peleška , Gustav Šír

Simulating Professions (SP) enables Large Language Models (LLMs) to emulate professional roles. However, comprehensive psychological and ethical evaluation in these contexts remains lacking. This paper introduces EMNLP, an Educator-role…

计算与语言 · 计算机科学 2025-12-19 Yilin Jiang , Mingzi Zhang , Sheng Jin , Zengyi Yu , Xiangjie Kong , Binghao Tu

Instruction-following benchmarks remain predominantly English-centric, leaving a critical evaluation gap for the hundreds of millions of Indic language speakers. We introduce IndicIFEval, a benchmark evaluating constrained generation of…

计算与语言 · 计算机科学 2026-02-26 Thanmay Jayakumar , Mohammed Safi Ur Rahman Khan , Raj Dabre , Ratish Puduppully , Anoop Kunchukuttan

The rapid adoption of generative artificial intelligence (AI) in educational assessment has created new opportunities for scalable item creation, personalized feedback, and efficient formative evaluation. However, despite advances in…

计算机与社会 · 计算机科学 2026-04-14 Antoun Yaacoub , Zainab Assaghir , Anuradha Kar

Large language models (LLMs) are increasingly used in K-12 education, yet existing benchmarks such as C-Eval, CMMLU, GaokaoBench, and EduEval mainly evaluate factual recall through exam-style question answering. Effective educational AI…

计算与语言 · 计算机科学 2026-05-12 Hao Liang , Qihan Lin , Zhaoyang Han , Xiaochen Ma , Zhen Hao Wong , Meiyi Qiang , Linzhuang Sun , Wentao Zhang

This study investigates how K-12 educators use generative AI tools in real-world instructional contexts and how large language models (LLMs) can support scalable qualitative analysis of these interactions. Drawing on over 13,000 unscripted…

人机交互 · 计算机科学 2025-12-17 Alex Liu , Lief Esbenshade , Shawon Sarkar , Victor Tian , Zachary Zhang , Kevin He , Min Sun

Advanced Large Multimodal Models (LMMs) have demonstrated impressive performance in K-12 reasoning tasks, exhibiting great promise as intelligent tutors. Realizing this potential requires models to navigate real-world examinations…

人工智能 · 计算机科学 2026-05-27 Xiaohan Wang , Mingze Yin , Yilin Zhao , Gang Liu , Dian Li