English
Related papers

Related papers: Assessing the Accuracy of Multisource Register-bas…

200 papers

We present an automatic piano transcription system that converts polyphonic audio recordings into musical scores. This has been a long-standing problem of music information processing, and recent studies have made remarkable progress in the…

Sound · Computer Science 2021-04-06 Kentaro Shibata , Eita Nakamura , Kazuyoshi Yoshii

Missing data is a ubiquitous problem. It is especially challenging in medical settings because many streams of measurements are collected at different - and often irregular - times. Accurate estimation of those missing measurements is…

Machine Learning · Computer Science 2017-11-27 Jinsung Yoon , William R. Zame , Mihaela van der Schaar

The Big Data era features a huge amount of data that are contributed by numerous sources and used by many critical data-driven applications. Due to the varying reliability of sources, it is common to see conflicts among the multi-source…

Databases · Computer Science 2017-08-08 Xiu Susie Fang , Quan Z. Sheng , Xianzhi Wang , Anne H. H. Ngu

Recent statistical evaluations for High-Energy Physics measurements, in particular those at the Large Hadron Collider, require careful evaluation of many sources of systematic uncertainties at the same time. While the fundamental aspects of…

Data Analysis, Statistics and Probability · Physics 2018-10-29 Luca Lista , Agostino De Iorio , Alberto Orso Maria Iorio

Food authenticity studies are concerned with determining if food samples have been correctly labeled or not. Discriminant analysis methods are an integral part of the methodology for food authentication. Motivated by food authenticity…

Methodology · Statistics 2010-10-08 Thomas Brendan Murphy , Nema Dean , Adrian E. Raftery

National statistical institutes currently investigate how to improve the output quality of official statistics based on machine learning algorithms. A key obstacle is concept drift, i.e., when the joint distribution of independent variables…

Methodology · Statistics 2021-03-02 Quinten Meertens , Cees Diks , Jaap van den Herik , Frank Takes

To improve the precision of inferences and reduce costs there is considerable interest in combining data from several sources such as sample surveys and administrative data. Appropriate methodology is required to ensure satisfactory…

Methodology · Statistics 2022-10-21 Dexter Cahoy , Joseph Sedransk

When averages of different experimental determinations of the same quantity are computed, each with statistical and systematic error components, then frequently the statistical and systematic components of the combined error are quoted…

Data Analysis, Statistics and Probability · Physics 2015-10-28 Jens Erler

Supervised machine learning assumes that labeled data provide accurate measurements of the concepts models are meant to learn. Yet in practice, human labeling introduces systematic variation arising from ambiguous items, divergent…

Methodology · Statistics 2026-04-10 Robert Chew , Stephanie Eckman , Christoph Kern , Frauke Kreuter

We examine the sources of error in the histogram reweighting method for Monte Carlo data analysis. We demonstrate that, in addition to the standard statistical error which has been studied elsewhere, there are two other sources of error,…

Statistical Mechanics · Physics 2015-06-25 M. E. J. Newman , R. G. Palmer

University evaluation is a topic of increasing concern in Italy as well as in other countries. In empirical analysis, university activities and performances are generally measured by means of indicator variables, summarizing the available…

Applications · Statistics 2014-04-25 Valentina Raponi , Francesca Martella , Antonello Maruotti

Many machine learning systems today are trained on large amounts of human-annotated data. Data annotation tasks that require a high level of competency make data acquisition expensive, while the resulting labels are often subjective,…

Machine Learning · Computer Science 2020-04-08 Emmanouil Antonios Platanios , Maruan Al-Shedivat , Eric Xing , Tom Mitchell

Return-to-baseline is an important method to impute missing values or unobserved potential outcomes when certain hypothetical strategies are used to handle intercurrent events in clinical trials. Current return-to-baseline approaches seen…

Methodology · Statistics 2021-11-19 Yongming Qu , Biyue Dai

A popular approach for large scale data annotation tasks is crowdsourcing, wherein each data point is labeled by multiple noisy annotators. We consider the problem of inferring ground truth from noisy ordinal labels obtained from multiple…

Machine Learning · Statistics 2013-05-02 Balaji Lakshminarayanan , Yee Whye Teh

This paper studies the change point problem for a general parametric, univariate or multivariate family of distributions. An information theoretic procedure is developed which is based on general divergence measures for testing the…

Statistics Theory · Mathematics 2014-03-26 Apostolos Batsidis , Nirian Martín , Leandro Pardo , Konstantinos Zografos

In over-identified models, misspecification -- the norm rather than exception -- fundamentally changes what estimators estimate. Different estimators imply different estimands rather than different efficiency for the same target. A review…

Econometrics · Economics 2026-02-23 Isaiah Andrews , Jiafeng Chen , Otavio Tecchio

Missing data is a fundamental challenge in data science, significantly hindering analysis and decision-making across a wide range of disciplines, including healthcare, bioinformatics, social science, e-commerce, and industrial monitoring.…

Machine Learning · Statistics 2026-05-12 Jicong Fan

In this report, we investigate (element-based) inconsistency measures for multisets of business rule bases. Currently, related works allow to assess individual rule bases, however, as companies might encounter thousands of such instances…

Artificial Intelligence · Computer Science 2021-03-02 Carl Corea , Matthias Thimm , Patrick Delfmann

Modern artificial intelligence is supported by machine learning models (e.g., foundation models) that are pretrained on a massive data corpus and then adapted to solve a variety of downstream tasks. To summarize performance across multiple…

Machine Learning · Statistics 2025-01-09 Rachel Longjohn , Giri Gopalan , Emily Casleton

The study presents a time-series analysis of field-standardized average impact of Italian research compared to the world average. The approach is purely bibliometric, based on census of the full scientific production from all Italian public…

Digital Libraries · Computer Science 2018-11-06 Giovanni Abramo , Ciriaco Andrea D'Angelo , Fulvio Viel