English
Related papers

Related papers: Online test administration results in students sel…

200 papers

We discuss multi-task online learning when a decision maker has to deal simultaneously with M tasks. The tasks are related, which is modeled by imposing that the M-tuple of actions taken by the decision maker needs to satisfy certain…

Machine Learning · Statistics 2009-03-27 Gabor Lugosi , Omiros Papaspiliopoulos , Gilles Stoltz

As large language models (LLMs) have grown in prevalence, particular benchmarks have become essential for the evaluation of these models and for understanding model capabilities. Most commonly, we use test accuracy averaged across multiple…

Computation and Language · Computer Science 2024-11-12 Vipul Gupta , David Pantoja , Candace Ross , Adina Williams , Megan Ung

Content-focused research-based assessment instruments typically use items (i.e., questions) as the unit of assessment for scoring, reporting, and validation. Couplet scoring employs an alternative unit of assessment called a couplet, which…

Physics Education · Physics 2024-12-11 Michael Vignal , Gayle Geschwind , Marcos D. Caballero , H. J. Lewandowski

When a Commit Message Generation (CMG) system is integrated into the IDEs and other products at JetBrains, we perform online evaluation based on user acceptance of the generated messages. However, performing online experiments with every…

Software Engineering · Computer Science 2025-01-09 Petr Tsvetkov , Aleksandra Eliseeva , Danny Dig , Alexander Bezzubov , Yaroslav Golubev , Timofey Bryksin , Yaroslav Zharov

Online controlled experiments, colloquially known as A/B-tests, are the bread and butter of real-world recommender system evaluation. Typically, end-users are randomly assigned some system variant, and a plethora of metrics are then…

Information Retrieval · Computer Science 2024-07-31 Olivier Jeunen , Shubham Baweja , Neeti Pokharna , Aleksei Ustimenko

IRT models are being increasingly used worldwide for test construction and scoring. The study examines the practical implications of estimating individual scores in a paper-and-pencil high-stakes test using 2PL and 3PL models, specifically…

Applications · Statistics 2018-05-03 Nancy Lacourly , Jaime San Martin , Monica Silva , Paula Uribe

Recommender systems are central to modern online platforms, but a popular concern is that they may be pulling society in dangerous directions (e.g., towards filter bubbles). However, a challenge with measuring the effects of recommender…

Computers and Society · Computer Science 2021-10-25 Serina Chang , Johan Ugander

Voting online with explicit ratings could largely reflect people's preferences and objects' qualities, but ratings are always irrational, because they may be affected by many unpredictable factors like mood, weather, as well as other…

Data Analysis, Statistics and Probability · Physics 2013-05-03 Zimo Yang , Zi-Ke Zhang , Tao Zhou

Large language models(LLMs) like Gemini are becoming common tools for supporting student writing. But most of their feedback is based only on the final essay missing important context about how that text was written. In this paper, we…

Computation and Language · Computer Science 2025-06-11 Samra Zafar , Shaheer Minhas , Syed Ali Hassan Zaidi , Arfa Naeem , Zahra Ali

Online grading systems have become extremely prevalent as majority of academic materials are in the process of being digitized, if not already done. In this paper, we present the concept of design and implementation of a mobile application…

Computers and Society · Computer Science 2023-11-13 Kaustubh Kundu , Sushant Yadav , Tayyabbali Sayyad

Click-through rate (CTR) prediction is a crucial task in online advertising to recommend products that users are likely to be interested in. To identify the best-performing models, rigorous model evaluation is necessary. Offline…

Information Retrieval · Computer Science 2024-06-27 Ramazan Tarik Turksoy , Beyza Turkmen

Learning an ordering of items based on pairwise comparisons is useful when items are difficult to rate consistently on an absolute scale, for example, when annotators have to make subjective assessments. When exhaustive comparison is…

Machine Learning · Computer Science 2024-10-29 Herman Bergström , Emil Carlsson , Devdatt Dubhashi , Fredrik D. Johansson

This paper discusses digital online mathematics examinations -- a discussion ranging from high school to university level examinations. In particular, we consider the nature of mathematical writing, what is distinctive about mathematical…

History and Overview · Mathematics 2026-05-26 Laura Kobel-Keller , Chris Sangwin

Item Response Theory (IRT) models aim to assess latent abilities of $n$ examinees along with latent difficulty characteristics of $m$ test items from categorical data that indicates the quality of their corresponding answers. Classical…

Machine Learning · Computer Science 2024-08-16 Susanne Frick , Amer Krivošija , Alexander Munteanu

Rubrics are being used in a wide variety of disciplines in higher education to evaluate assessments and provide feedback to students. Rubrics are traditionally implemented as paper-based table format to grade assessments and provide…

Computers and Society · Computer Science 2016-06-07 Phil Smith , Mohan John Blooma , Jayan Kurian

We propose a simple refactoring of multi-choice question answering (MCQA) tasks as a series of binary classifications. The MCQA task is generally performed by scoring each (question, answer) pair normalized over all the pairs, and then…

Computation and Language · Computer Science 2022-11-01 Deepanway Ghosal , Navonil Majumder , Rada Mihalcea , Soujanya Poria

Feedback is a critical component of the learning process, particularly in computer science education. This study investigates the quality of feedback generated by Large Language Models (LLMs), Small Language Models (SLMs), compared with…

Human-Computer Interaction · Computer Science 2026-01-21 Suqing Liu , Bogdan Simion , Christopher Eaton , Michael Liut

Ishimoto, Davenport, and Wittmann have previously reported analyses of data from student responses to the Force and Motion Conceptual Evaluation (FMCE), in which they used item response curves (IRCs) to make claims about American and…

Physics Education · Physics 2021-10-04 Connor J. Richardson , Trevor I. Smith , Paul J. Walter

Recent years have seen enormous gains in core IR tasks, including document and passage ranking. Datasets and leaderboards, and in particular the MS MARCO datasets, illustrate the dramatic improvements achieved by modern neural rankers. When…

Information Retrieval · Computer Science 2022-03-02 Negar Arabzadeh , Alexandra Vtyurina , Xinyi Yan , Charles L. A. Clarke

Mobile learning (m-Learning) is considered to be one of the fastest growing learning platforms. The immense interest in m-Learning is attributed to the incredible rate of growth of mobile technology and its proliferation into every aspect…

Computers and Society · Computer Science 2018-01-15 Muasaad Alrasheedi , Luiz Fernando Capretz , Arif Raza
‹ Prev 1 8 9 10 Next ›