English
Related papers

Related papers: On Contamination of Symbolic Datasets

200 papers

Data contamination in model evaluation has become increasingly prevalent with the growing popularity of large language models. It allows models to "cheat" via memorisation instead of displaying true capabilities. Therefore, contamination…

Computation and Language · Computer Science 2024-01-30 Yucheng Li , Frank Guerin , Chenghua Lin

Recent studies apply psychometric questionnaires to Large Language Models (LLMs) to assess high-level psychological constructs such as values, personality, moral foundations, and dark traits. Although prior work has raised concerns about…

Computation and Language · Computer Science 2026-02-02 Jongwook Han , Woojung Song , Jonggeun Lee , Yohan Jo

Identifying anomalies and contamination in datasets is important in a wide variety of settings. In this paper, we describe a new technique for estimating contamination in large, discrete valued datasets. Our approach considers the normal…

Information Theory · Computer Science 2015-06-16 Matthew L. Malloy , Scott Alfeld , Paul Barford

High throughput metabolomics data are fraught with both non-ignorable missing observations and unobserved factors that influence a metabolite's measured concentration, and it is well known that ignoring either of these complications can…

Methodology · Statistics 2019-09-09 Chris McKennan , Carole Ober , Dan Nicolae

The detectors in mass spectrometers are precise enough to count ion events. In practice, the statistics of chemical noise are affected by large quantization errors and overdispersion because of amplification in the detector. The detector…

Discrete Mathematics · Computer Science 2009-06-02 Sébastien Li-Thiao-Té

When an individual's DNA is sequenced, sensitive medical information becomes available to the sequencing laboratory. A recently proposed way to hide an individual's genetic information is to mix in DNA samples of other individuals. We…

Information Theory · Computer Science 2024-11-05 Kayvon Mazooji , Roy Dong , Ilan Shomorony

Machine learning (ML) datasets, often perceived as neutral, inherently encapsulate abstract and disputed social constructs. Dataset curators frequently employ value-laden terms such as diversity, bias, and quality to characterize datasets.…

Machine Learning · Computer Science 2024-07-12 Dora Zhao , Jerone T. A. Andrews , Orestis Papakyriakopoulos , Alice Xiang

Machine learning models encounter Out-of-Distribution (OoD) errors when the data seen at test time are generated from a different stochastic generator than the one used to generate the training data. One proposal to scale OoD detection to…

Machine Learning · Statistics 2019-05-27 Hyunsun Choi , Eric Jang , Alexander A. Alemi

In this paper, we study a classification problem in which sample labels are randomly corrupted. In this scenario, there is an unobservable sample with noise-free labels. However, before being observed, the true labels are independently…

Machine Learning · Statistics 2015-07-21 Tongliang Liu , Dacheng Tao

In many studies of human diseases, multiple omic datasets are measured. Typically, these omic datasets are studied one by one with the disease, thus the relationship between omics are overlooked. Modeling the joint part of multiple omics…

Methodology · Statistics 2022-09-02 Zhujie Gu , Said el Bouhaddani , Jeanine Houwing-Duistermaat , Hae-Won Uh

Modeling real-world systems requires accounting for noise - whether it arises from unpredictable fluctuations in financial markets, irregular rhythms in biological systems, or environmental variability in ecosystems. While the behavior of…

Machine Learning · Computer Science 2026-04-08 Matteo Bosso , Giovanni Franzese , Kushal Swamy , Maarten Theulings , Alejandro M. Aragón , Farbod Alijani

Consider making a prediction over new test data without any opportunity to learn from a training set of labelled data - instead given access to a set of expert models and their predictions alongside some limited information about the…

Machine Learning · Computer Science 2022-10-12 Alex J. Chan , Mihaela van der Schaar

We characterise the evolution of a dynamical system by combining two well-known complex systems' tools, namely, symbolic ordinal analysis and networks. From the ordinal representation of a time-series we construct a network in which every…

Background: The analysis of DNA methylation is a key component in the development of personalized treatment approaches. A common way to measure DNA methylation is the calculation of beta values, which are bounded variables of the form M =…

Methodology · Statistics 2016-07-26 Leonie Weinhold , Simone Wahl , Matthias Schmid

Since the behavior of a neural network model is adversely affected by a lack of diversity in training data, we present a method that identifies and explains such deficiencies. When a dataset is labeled, we note that annotations alone are…

Computer Vision and Pattern Recognition · Computer Science 2020-12-17 Dhasarathy Parthasarathy , Anton Johansson

A finite set is "hidden" if its elements are not directly enumerable or if its size cannot be ascertained via a deterministic query. In public health, epidemiology, demography, ecology and intelligence analysis, researchers have developed a…

Statistics Theory · Mathematics 2019-10-17 Si Cheng , Daniel J. Eck , Forrest W. Crawford

In embodied intelligence, datasets play a pivotal role, serving as both a knowledge repository and a conduit for information transfer. The two most critical attributes of a dataset are the amount of information it provides and how easily…

Robotics · Computer Science 2025-11-13 Jiahao Xiao , Bowen Yan , Jianbo Zhang , Jia Wang , Chunyi Li , Zhengxue Cheng , Guangtao Zhai

Memorization in over-parameterized neural networks could severely hurt generalization in the presence of mislabeled examples. However, mislabeled examples are hard to avoid in extremely large datasets collected with weak supervision. We…

Machine Learning · Computer Science 2020-04-10 Jiaming Song , Lunjia Hu , Michael Auli , Yann Dauphin , Tengyu Ma

We developed a single factor model with measure-specific sample weights for multivariate data with multiple observed indicators clustered within a higher level subject. The factor is therefore a latent variable shared by multiple indicators…

Methodology · Statistics 2019-10-22 Chengan Du , Shu-Xia Li , Zhenqiu Lin , Haiqun Lin

Deep neural networks can memorize corrupted labels, making data quality critical for model performance, yet real-world datasets are frequently compromised by both label noise and input noise. This paper proposes a mutual information-based…

Machine Learning · Computer Science 2025-08-12 Jinghan Yang , Jiayu Weng
‹ Prev 1 3 4 5 6 7 10 Next ›