English
Related papers

Related papers: On Data-centric Myths

200 papers

Today, huge amounts of data are being collected with spatial and temporal components from sources such as meteorological, satellite imagery etc. Efficient visualisation as well as discovery of useful knowledge from these datasets is…

Databases · Computer Science 2017-03-31 Nhien-An Le-Khac , Martin Bue , Michael Whelan , Tahar Kechadi

Data augmentation has been widely applied as an effective methodology to improve generalization in particular when training deep neural networks. Recently, researchers proposed a few intensive data augmentation techniques, which indeed…

Machine Learning · Computer Science 2019-11-22 Zhuoxun He , Lingxi Xie , Xin Chen , Ya Zhang , Yanfeng Wang , Qi Tian

Association rule mining is an important data-mining technique that finds interesting association among a large set of data items. Since it may disclose patterns and various kinds of sensitive knowledge that are difficult to find otherwise,…

Databases · Computer Science 2012-04-10 Dhyanendra Jain

The recent interest in Big Data has generated a broad range of new academic, corporate, and policy practices along with an evolving debate amongst its proponents, detractors, and skeptics. While the practices draw on a common set of tools,…

Bayesian inference for inverse problems hinges critically on the choice of priors. In the absence of specific prior information, population-level distributions can serve as effective priors for parameters of interest. With the advent of…

Instrumentation and Methods for Astrophysics · Physics 2025-02-11 Gabriel Missael Barco , Alexandre Adam , Connor Stone , Yashar Hezaveh , Laurence Perreault-Levasseur

Distribution shift is a major source of failure for machine learning models. However, evaluating model reliability under distribution shift can be challenging, especially since it may be difficult to acquire counterfactual examples that…

Machine Learning · Computer Science 2023-06-21 Joshua Vendrow , Saachi Jain , Logan Engstrom , Aleksander Madry

Motivated by recently emerging problems in machine learning and statistics, we propose data models which relax the familiar i.i.d. assumption. In essence, we seek to understand what it means for data to come from a set of probability…

Statistics Theory · Mathematics 2025-01-08 Christian Fröhlich , Robert C. Williamson

Identifying model parameters from observed configurations poses a fundamental challenge in data science, especially with limited data. Recently, diffusion models have emerged as a novel paradigm in generative machine learning, capable of…

Data Analysis, Statistics and Probability · Physics 2025-03-14 Yechan Lim , Sangwon Lee , Junghyo Jo

We propose a rigorous decomposition of predictive error, highlighting that not all 'irreducible' error is genuinely immutable. Many domains stand to benefit from iterative enhancements in measurement, construct validity, and modeling. Our…

Machine Learning · Computer Science 2025-02-12 Jiani Yan , Charles Rahal

The goal of contrasting learning is to learn a representation that preserves underlying clusters by keeping samples with similar content, e.g. the ``dogness'' of a dog, close to each other in the space generated by the representation. A…

Machine Learning · Computer Science 2023-02-17 Advait Parulekar , Liam Collins , Karthikeyan Shanmugam , Aryan Mokhtari , Sanjay Shakkottai

A key element in transfer learning is representation learning; if representations can be developed that expose the relevant factors underlying the data, then new tasks and domains can be learned readily based on mappings of these salient…

Machine Learning · Computer Science 2014-12-18 Yujia Li , Kevin Swersky , Richard Zemel

We study large-scale classification problems in changing environments where a small part of the dataset is modified, and the effect of the data modification must be quickly incorporated into the classifier. When the entire dataset is large,…

Machine Learning · Statistics 2016-06-02 Hiroyuki Hanada , Atsushi Shibagaki , Jun Sakuma , Ichiro Takeuchi

In this work we consider the problem of data classification in post-classical settings were the number of training examples consists of mere few data points. We explore the phenomenon and reveal key relationships between dimensionality of…

Machine Learning · Computer Science 2022-04-01 Ivan Y. Tyukin , Oliver Sutton , Alexander N. Gorban

The increasing availability of passively observed data has yielded a growing methodological interest in "data fusion." These methods involve merging data from observational and experimental sources to draw causal conclusions -- and they…

Methodology · Statistics 2021-12-15 Evan Rosenman , Art B. Owen

Big data analysis poses the dual problem of privacy preservation and utility, i.e., how accurate data analyses remain after transforming original data in order to protect the privacy of the individuals that the data is about - and whether…

Machine Learning · Computer Science 2022-11-29 Md Sakib Nizam Khan , Niklas Reje , Sonja Buchegger

Data representativity is crucial when drawing inference from data through machine learning models. Scholars have increased focus on unraveling the bias and fairness in models, also in relation to inherent biases in the input data. However,…

Machine Learning · Statistics 2023-02-06 Line H. Clemmensen , Rune D. Kjærsgaard

The central theme of this talk is to promote the non-asymptotic statistical viewpoint in the context of massive datasets. The classical viewpoint breaks down when the data size becomes large.

Information Theory · Computer Science 2014-12-23 Robert C. Qiu

Data-driven predictive solutions predominant in commercial applications tend to suffer from biases and stereotypes, which raises equity concerns. Prediction models may discover, use, or amplify spurious correlations based on gender or other…

Computation and Language · Computer Science 2022-11-28 Abdelrahman Zayed , Prasanna Parthasarathi , Goncalo Mordido , Hamid Palangi , Samira Shabanian , Sarath Chandar

Experimental life sciences like biology or chemistry have seen in the recent decades an explosion of the data available from experiments. Laboratory instruments become more and more complex and report hundreds or thousands measurements for…

Machine Learning · Statistics 2014-03-13 C. O. S. Sorzano , J. Vargas , A. Pascual Montano

Data Science is currently a popular field of science attracting expertise from very diverse backgrounds. Current learning practices need to acknowledge this and adapt to it. This paper summarises some experiences relating to such learning…

General Literature · Computer Science 2018-07-11 Yehia Elkhatib