中文
相关论文

相关论文: Rasch Analysis of the Mathematics Self Concept Que…

200 篇论文

The Fields Medal, often referred as the Nobel Prize of mathematics, is awarded to no more than four mathematician under the age of 40, every four years. In recent years, its conferral has come under scrutiny of math historians, for…

历史与综述 · 数学 2020-02-19 Ho-Chun Herbert Chang , Feng Fu

Classroom discourse is a core medium of instruction - analyzing it can provide a window into teaching and learning as well as driving the development of new tools for improving instruction. We introduce the largest dataset of mathematics…

计算与语言 · 计算机科学 2023-05-26 Dorottya Demszky , Heather Hill

This study discusses an alternative tool for modeling student assessment data. The model constructs networks from a matrix item responses and attempts to represent these data in low dimensional Euclidean space. This procedure has advantages…

应用统计 · 统计学 2020-03-18 Alex Brodersen , Ick Hoon Jin , Ying Cheng , Minjeong Jeon

Self-evolving reasoning frameworks let LLMs improve their reasoning capabilities by iteratively generating and solving problems without external supervision, using verifiable rewards. Ideally, such systems are expected to explore a diverse…

机器学习 · 计算机科学 2026-03-17 Vaibhav Mishra

The International Symposium on Computational Sensing (ISCS) brings together researchers from optical microscopy, electron microscopy, RADAR, astronomical imaging, biomedical imaging, remote sensing, and signal processing. With a particular…

信号处理 · 电气工程与系统科学 2023-08-30 Thomas Feuillen , Amirafshar Moshtaghpour

Math anxiety is a clinical pathology impairing cognitive processing in math-related contexts. Originally thought to affect only inexperienced, low-achieving students, recent investigations show how math anxiety is vastly diffused even among…

社会与信息网络 · 计算机科学 2021-09-01 Massimo Stella

To bridge the gap between performance-oriented benchmarks and the evaluation of cognitively inspired models, we introduce BLiSS 1.0, a Benchmark of Learner Interlingual Syntactic Structure. Our benchmark operationalizes a new paradigm of…

计算与语言 · 计算机科学 2025-10-23 Yuan Gao , Suchir Salhan , Andrew Caines , Paula Buttery , Weiwei Sun

This study investigates how students and researchers shape their knowledge and perception of educational topics. The mindset or forma mentis of 159 Italian high school students and of 59 international researchers in STEM are reconstructed…

物理教育 · 物理学 2020-01-09 Massimo Stella

LLMs have demonstrated remarkable capability for understanding semantics, but they often struggle with understanding pragmatics. To demonstrate this fact, we release a Pragmatics Understanding Benchmark (PUB) dataset consisting of fourteen…

This paper introduces a new spreadsheet tool for adoption by high school or college level physics teachers who use common assessments in a pre-instruction/post-instruction mode to diagnose student learning and teaching effectiveness. The…

物理教育 · 物理学 2017-03-14 Gary A. Morris , Paul J. Walter , Spencer Skees , Samantha Schwartz

Temporal reasoning is pivotal for Large Language Models (LLMs) to comprehend the real world. However, existing works neglect the real-world challenges for temporal reasoning: (1) intensive temporal information, (2) fast-changing event…

人工智能 · 计算机科学 2025-10-09 Shaohang Wei , Wei Li , Feifan Song , Wen Luo , Tianyi Zhuang , Haochen Tan , Zhijiang Guo , Houfeng Wang

The evaluation of instructors by their students has been practiced at most universities for many decades, and there has always been a great interest in a variety of aspects of the evaluations. Are students matured and knowledgeable enough…

应用统计 · 统计学 2015-01-12 Necla Gunduz , Ernest Fokoue

As the mathematical capabilities of large language models (LLMs) improve, it becomes increasingly important to evaluate their performance on research-level tasks at the frontier of mathematical knowledge. However, existing benchmarks are…

Evaluation of Text to Speech (TTS) systems is challenging and resource-intensive. Subjective metrics such as Mean Opinion Score (MOS) are not easily comparable between works. Objective metrics are frequently used, but rarely validated…

声音 · 计算机科学 2026-03-03 Christoph Minixhofer , Ondrej Klejch , Peter Bell

Proficiency with calculating, reporting, and understanding measurement uncertainty is a nationally recognized learning outcome for undergraduate physics lab courses. The Physics Measurement Questionnaire (PMQ) is a research-based assessment…

Traditional problem-based exams are not efficient instruments for assessing the "structure" of physics students' conceptual knowledge or for providing diagnostically detailed feedback to students and teachers. We present the Free Term Entry…

物理教育 · 物理学 2007-05-23 Ian D. Beatty , William J. Gerace , Robert J. Dufresne

The rapid advancement of large language models has opened new avenues for automating complex problem-solving tasks such as algorithmic coding and competitive programming. This paper introduces a novel evaluation technique, LLM-ProS, to…

计算与语言 · 计算机科学 2026-03-03 Md Sifat Hossain , Anika Tabassum , Md. Fahim Arefin , Tarannum Shaila Zaman

With the increasing integration of large lauguage models (LLMs) in education, there is growing interest in using AI agents to support student learning in creative tasks. This study presents an interactive Mentor Agent system named Mentigo,…

人机交互 · 计算机科学 2024-09-24 Siyu Zha , Yujia Liu , Chengbo Zheng , Jiaqi XU , Fuze Yu , Jiangtao Gong , Yingqing XU

The rapid growth in submissions to machine learning venues has strained the scientific peer-review system and intensified interest in LLM-based automated peer reviewers. However, how good these systems are actually, especially compared to…

We present a test corpus of audio recordings and transcriptions of presentations of students' enterprises together with their slides and web-pages. The corpus is intended for evaluation of automatic speech recognition (ASR) systems,…

计算与语言 · 计算机科学 2019-08-05 Dominik Macháček , Jonáš Kratochvíl , Tereza Vojtěchová , Ondřej Bojar