中文
相关论文

相关论文: On Evaluation of Vision Datasets and Models using …

200 篇论文

Educational assessment relies heavily on knowing question difficulty, traditionally determined through resource-intensive pre-testing with students. This creates significant barriers for both classroom teachers and assessment developers. We…

计算机与社会 · 计算机科学 2026-02-03 Matias Hoyl

Multidimensional item response theory is a statistical test theory used to estimate the latent skills of learners and the difficulty levels of problems based on test results. Both compensatory and non-compensatory models have been proposed…

统计方法学 · 统计学 2025-07-22 Hiroshi Tamano , Hideitsu Hino , Daichi Mochihashi

Item parameter estimation in pharmacometric item response theory (IRT) models is predominantly performed using the Laplace estimation algorithm as implemented in NONMEM. In psychometrics a wide range of different software tools, including…

统计方法学 · 统计学 2025-03-18 Leticia Arrington , Sebastian Ueckert

Analyses of heterogeneous treatment effects (HTE) are common in applied causal inference research. However, when outcomes are latent variables assessed via psychometric instruments such as educational tests, standard methods ignore the…

计量经济学 · 经济学 2025-06-27 Joshua B. Gilbert , Zachary Himmelsbach , James Soland , Mridul Joshi , Benjamin W. Domingue

Item Response Theory (IRT) was originally developed in traditional exam settings, and it has been shown that the model does not readily transfer to formative assessment in the form of online homework. We investigate if this is mostly due to…

物理教育 · 物理学 2015-03-24 Emre Gönülateş , Gerd Kortemeyer

Assessment of proficiency of the learner is an essential part of Intelligent Tutoring Systems (ITS). We use Item Response Theory (IRT) in computer-aided language learning for assessment of student ability in two contexts: in test sessions,…

人工智能 · 计算机科学 2024-09-25 Jue Hou , Anisia Katinskaia , Anh-Duc Vu , Roman Yangarber

Large language models (LLMs) have demonstrated exceptional performance across a wide range of natural language tasks. However, selecting the optimal LLM to respond to a user query often necessitates a delicate balance between performance…

人工智能 · 计算机科学 2025-06-24 Wei Song , Zhenya Huang , Cheng Cheng , Weibo Gao , Bihan Xu , GuanHao Zhao , Fei Wang , Runze Wu

Most datasets suffer from partial or complete missing values, which has downstream limitations on the available models on which to test the data and on any statistical inferences that can be made from the data. Several imputation techniques…

机器学习 · 统计学 2023-02-09 Adrienne Kline , Yuan Luo

Since early machine learning models, metrics such as accuracy and precision have been the de facto way to evaluate and compare trained models. However, a single metric number doesn't fully capture the similarities and differences between…

计算机视觉与模式识别 · 计算机科学 2022-08-02 Ahmad Mustapha , Wael Khreich , Wes Masri

Item difficulty plays a crucial role in test performance, interpretability of scores, and equity for all test-takers, especially in large-scale assessments. Traditional approaches to item difficulty modeling rely on field testing and…

计算与语言 · 计算机科学 2025-09-30 Sydney Peters , Nan Zhang , Hong Jiao , Ming Li , Tianyi Zhou , Robert Lissitz

Item response theory (IRT) models are a class of statistical models used to describe the response behaviors of individuals to a set of items having a certain number of options. They are adopted by researchers in social science, particularly…

统计计算 · 统计学 2014-04-16 Angelo Mazza , Antonio Punzo , Brian McGuire

Deep learning based knowledge tracing model has been shown to outperform traditional knowledge tracing model without the need for human-engineered features, yet its parameters and representations have long been criticized for not being…

机器学习 · 计算机科学 2019-04-29 Chun-Kit Yeung

Item response theory (IRT) is a non-linear generative probabilistic paradigm for using exams to identify, quantify, and compare latent traits of individuals, relative to their peers, within a population of interest. In pre-existing…

机器学习 · 计算机科学 2019-12-06 Joshua C. Chang , Shashaank Vattikuti , Carson C. Chow

Despite the availability of benchmark machine learning (ML) repositories (e.g., UCI, OpenML), there is no standard evaluation strategy yet capable of pointing out which is the best set of datasets to serve as gold standard to test different…

In this paper, we apply Item Response Theory, popular in education and political science research, to the analysis of argument persuasiveness in language. We empirically evaluate the model's performance on three datasets, including a novel…

计算与语言 · 计算机科学 2022-04-26 Anastassia Kornilova , Daniel Argyle , Vladimir Eidelman

While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offering limited insight into whether LLM judges themselves function as stable and reliable…

人工智能 · 计算机科学 2026-02-03 Junhyuk Choi , Sohhyung Park , Chanhee Cho , Hyeonchu Park , Bugeun Kim

Multimodal Large Language Models (MLLMs) have recently emerged as general architectures capable of reasoning over diverse modalities. Benchmarks for MLLMs should measure their ability for cross-modal integration. However, current benchmarks…

计算与语言 · 计算机科学 2026-03-04 Shunki Uebayashi , Kento Masui , Kyohei Atarashi , Han Bao , Hisashi Kashima , Naoto Inoue , Mayu Otani , Koh Takeuchi

We propose a class of Item Response Theory models for items with ordinal polytomous responses, which extends an existing class of multidimensional models for dichotomously-scored items measuring more than one latent trait. In the proposed…

统计方法学 · 统计学 2012-01-24 Silvia Bacci , Francesco Bartolucci , Michela Gnaldi

Machine learning has shown much promise in helping improve the quality of medical, legal, and financial decision-making. In these applications, machine learning models must satisfy two important criteria: (i) they must be causal, since the…

机器学习 · 计算机科学 2021-10-12 Carolyn Kim , Osbert Bastani

While Vision-Language Models (VLMs) have achieved competitive performance in various tasks, their comprehension of the underlying structure and semantics of a scene remains understudied. To investigate the understanding of VLMs, we study…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Massimo Rizzoli , Simone Alghisi , Olha Khomyn , Gabriel Roccabruna , Seyed Mahed Mousavi , Giuseppe Riccardi