English
Related papers

Related papers: A Supervised Hybrid Statistical Catch-up System Bu…

200 papers

We consider an evolving system for which a sequence of observations is being made, with each observation revealing additional information about current and past states of the system. We suppose each observation is made without error, but…

Computation · Statistics 2021-03-10 Valentina Di Marco , Jonathan Keith

We consider studies where multiple measures on an outcome variable are collected over time, but some subjects drop out before the end of follow up. Analyses of such data often proceed under either a 'last observation carried forward' or…

Methodology · Statistics 2022-07-26 Oliver Dukes , David Richardson , Eric Tchetgen Tchetgen

Double-descent refers to the unexpected drop in test loss of a learning algorithm beyond an interpolating threshold with over-parameterization, which is not predicted by information criteria in their classical forms due to the limitations…

Machine Learning · Computer Science 2023-11-15 Haobo Chen , Yuheng Bu , Gregory W. Wornell

We present a new method for jointly modelling the students' results in the university's admission exams and their performance in subsequent courses at the university. The case considered involved all the students enrolled at the University…

Applications · Statistics 2021-02-23 Jeanett S. Pelck , Rafael Pimentel Maia , Hildete P. Pinheiro , Rodrigo Labouriau

Motivated by real-world machine learning applications, we consider a statistical classification task in a sequential setting where test samples arrive sequentially. In addition, the generating distributions are unknown and only a set of…

Machine Learning · Statistics 2021-02-11 Mahdi Haghifam , Vincent Y. F. Tan , Ashish Khisti

Educational Data Mining (EDM) is a developing discipline, concerned with expanding the classical Data Mining (DM) methods and developing new methods for discovering the data that originate from educational systems. Student attendance in…

Computers and Society · Computer Science 2020-09-03 Mohammed Alsuwaiket , Christian Dawson , Firat Batmaz

In supervised learning, we often face with ambiguous (A) samples that are difficult to label even by domain experts. In this paper, we consider a binary classification problem in the presence of such A samples. This problem is substantially…

Machine Learning · Computer Science 2020-11-25 Naoya Otani , Yosuke Otsubo , Tetsuya Koike , Masashi Sugiyama

We investigate model based classification with partially labelled training data. In many biostatistical applications, labels are manually assigned by experts, who may leave some observations unlabelled due to class uncertainty. We analyse…

Methodology · Statistics 2019-04-08 Daniel Ahfock , Geoffrey J. McLachlan

The increased availability of observation data from engineering systems in operation poses the question of how to incorporate this data into finite element models. To this end, we propose a novel statistical construction of the finite…

Methodology · Statistics 2021-01-25 Mark Girolami , Eky Febrianto , Ge Yin , Fehmi Cirak

It has been argued for many years that models used to analyze data from crossover designs are not appropriate when simple carryover effects are assumed. Furthermore, a statistical model that could estimate complex carry-over effects in…

Methodology · Statistics 2025-08-22 N. A. Cruz , K. Mylona , O. O. Melo

The increased availability of data in recent years has led several authors to ask whether it is possible to use data as a {\em computational} resource. That is, if more data is available, beyond the sample complexity limit, is it possible…

Machine Learning · Computer Science 2013-11-12 Amit Daniely , Nati Linial , Shai Shalev Shwartz

In Bayesian statistics, the marginal likelihood, also known as the evidence, is used to evaluate model fit as it quantifies the joint probability of the data under the prior. In contrast, non-Bayesian models are typically compared using…

Methodology · Statistics 2019-09-24 Edwin Fong , Chris Holmes

Deep learning has revolutionized medical imaging, but its effectiveness is severely limited by insufficient labeled training data. This paper introduces a novel GAN-based semi-supervised learning framework specifically designed for low…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Guido Manni , Clemente Lauretti , Loredana Zollo , Paolo Soda

The primary objective of this report is to determine what influences the success rates of students who have studied in Colombia, analyzing the Saber 11, the test done at the last school year, some socioeconomic aspects and comparing the…

Computers and Society · Computer Science 2020-06-03 Gregorio Perez Bernal , Luisa Toro Villegas , Mauricio Toro

We study the role of information and access in capacity-constrained selection problems with fairness concerns. We develop a statistical discrimination framework, where each applicant has multiple features and is potentially strategic. The…

Computers and Society · Computer Science 2026-03-03 Nikhil Garg , Hannah Li , Faidra Monachou

U.S. state education agencies mark schools displaying achievement gaps between demographic subgroups as needing improvement. Some schools may have few students in these subgroups, such that average end-of-year test scores only noisily…

Methodology · Statistics 2025-12-10 Joshua Wasserman , Michael R. Elliott , Ben B. Hansen

The analysis of competing risks data is often complicated by misclassification of the cause of failure. This issue can lead to seriously biased estimates and invalid conclusions. One way to deal with such misclassification is to use a…

This study investigates the factors associated with failure in each of the four thematic units of a General Statistics course offered at a private university in Colombia. Unlike traditional analyses that treat performance as a single…

Other Statistics · Statistics 2025-10-24 Biviana Marcela Suarez Sierra

Three important issues are often encountered in Supervised and Semi-Supervised Classification: class-memberships are unreliable for some training units (label noise), a proportion of observations might depart from the main structure of the…

Applications · Statistics 2020-07-02 Andrea Cappozzo , Francesca Greselin , Thomas Brendan Murphy

Statisticians have recently developed propensity score methods to improve generalizations from randomized experiments that do not employ random sampling. However, these methods typically rely on assumptions whose plausibility may be…

Methodology · Statistics 2019-11-14 Wendy Chan