English
Related papers

Related papers: Towards Formalizing Spuriousness of Biased Dataset…

200 papers

The properties of complex networked systems arise from the interplay between the dynamics of their elements and the underlying topology. Thus, to understand their behaviour, it is crucial to convene as much information as possible about…

Neurons and Cognition · Quantitative Biology 2024-06-18 Gustavo Menesse , Akke Mats Houben , Jordi Soriano , Joaquin J. Torres

We introduce a two-stage probabilistic framework for statistical downscaling using unpaired data. Statistical downscaling seeks a probabilistic map to transform low-resolution data from a biased coarse-grained numerical scheme to…

Machine Learning · Computer Science 2023-11-01 Zhong Yi Wan , Ricardo Baptista , Yi-fan Chen , John Anderson , Anudhyan Boral , Fei Sha , Leonardo Zepeda-Núñez

The apparent dichotomy between information-processing and dynamical approaches to complexity science forces researchers to choose between two diverging sets of tools and explanations, creating conflict and often hindering scientific…

Neurons and Cognition · Quantitative Biology 2022-01-26 Pedro A. M. Mediano , Fernando E. Rosas , Juan Carlos Farah , Murray Shanahan , Daniel Bor , Adam B. Barrett

A key objective of decomposition analysis is to identify a factor (the 'mediator') contributing to disparities in an outcome between social groups. In decomposition analysis, a scholarly interest often centers on estimating how much the…

Methodology · Statistics 2022-05-27 Soojin Park , Suyeon Kang , Chioun Lee , Shujie Ma

The inference of causal relationships using observational data from partially observed multivariate systems with hidden variables is a fundamental question in many scientific domains. Methods extracting causal information from conditional…

Machine Learning · Statistics 2020-10-13 Daniel Chicharro , Michel Besserve , Stefano Panzeri

The performance of deep neural networks is strongly influenced by the training dataset setup. In particular, when attributes having a strong correlation with the target attribute are present, the trained model can provide unintended…

Machine Learning · Computer Science 2023-02-14 Sumyeong Ahn , Seongyoon Kim , Se-young Yun

Social scientists are increasingly turning to unstructured datasets to unlock new empirical insights, e.g., estimating descriptive statistics of or causal effects on quantitative measures derived from text, audio, or video data. In many…

Econometrics · Economics 2026-05-06 Jacob Carlson

The ability to learn disentangled representations that split underlying sources of variation in high dimensional, unstructured data is important for data efficient and robust use of neural networks. While various approaches aiming towards…

Machine Learning · Statistics 2019-05-15 Raphael Suter , Đorđe Miladinović , Bernhard Schölkopf , Stefan Bauer

Reliable estimation of predictive uncertainty is crucial for machine learning applications, particularly in high-stakes scenarios where hedging against risks is essential. Despite its significance, there is no universal agreement on how to…

Machine Learning · Computer Science 2025-06-17 Kajetan Schweighofer , Lukas Aichberger , Mykyta Ielanskyi , Sepp Hochreiter

Image enhancement approaches often assume that the noise is signal independent, and approximate the degradation model as zero-mean additive Gaussian. However, this assumption does not hold for biomedical imaging systems where sensor-based…

Image and Video Processing · Electrical Eng. & Systems 2023-04-10 Calvin-Khang Ta , Abhishek Aich , Akash Gupta , Amit K. Roy-Chowdhury

Spurious correlations are a major source of errors for machine learning models, in particular when aiming for group-level fairness. It has been recently shown that a powerful approach to combat spurious correlations is to re-train the last…

Machine Learning · Computer Science 2024-09-24 Humza Wajid Hameed , Geraldin Nanfack , Eugene Belilovsky

Aleatoric (data) and epistemic (knowledge) uncertainty are textbook components of Uncertainty Quantification. Jointly estimating these components has been shown to be problematic and non-trivial. As a result, there are multiple ways to…

Machine Learning · Computer Science 2026-02-12 Ivo Pascal de Jong , Andreea Ioana Sburlea , Matthia Sabatelli , Matias Valdenegro-Toro

Deep learning models can excel on medical tasks, yet often experience spurious correlations, known as shortcut learning, leading to poor generalization in new environments. Particularly in medical imaging, where multiple spurious…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Louisa Fay , Hajer Reguigui , Bin Yang , Sergios Gatidis , Thomas Küstner

The advancements in disentangled representation learning significantly enhance the accuracy of counterfactual predictions by granting precise control over instrumental variables, confounders, and adjustable variables. An appealing method…

Machine Learning · Computer Science 2024-06-17 Xinshu Li , Mingming Gong , Lina Yao

We consider biological individuality in terms of information theoretic and graphical principles. Our purpose is to extract through an algorithmic decomposition system-environment boundaries supporting individuality. We infer or detect…

Populations and Evolution · Quantitative Biology 2014-12-09 David Krakauer , Nils Bertschinger , Eckehard Olbrich , Nihat Ay , Jessica C. Flack

A fundamental task in science is to determine the underlying causal relations because it is the knowledge of this functional structure what leads to the correct interpretation of an effect given the apparent associations in the observed…

Artificial Intelligence · Computer Science 2024-08-02 Alexandre Trilla , Nenad Mijatovic

Quantifying modality contributions in multimodal models remains a challenge, as existing approaches conflate the notion of contribution itself. Prior work relies on accuracy-based approaches, interpreting performance drops after removing a…

Machine Learning · Computer Science 2025-11-26 Padegal Amit , Omkar Mahesh Kashyap , Namitha Rayasam , Nidhi Shekhar , Surabhi Narayan

We characterize information as risk reduction between knowledge states represented by partitions of the underlying probability space. Entropy corresponds to risk reduction from no (or partial) knowledge to full knowledge about a random…

Information Theory · Computer Science 2026-02-24 Sebastian Gottwald , Daniel A. Braun

A popular method for variance reduction in observational causal inference is propensity-based trimming, the practice of removing units with extreme propensities from the sample. This practice has theoretical grounding when the data are…

Methodology · Statistics 2024-01-30 Samir Khan , Johan Ugander

Identifying and disentangling sources of predictive uncertainty is essential for trustworthy supervised learning. We argue that widely used second-order methods that disentangle aleatoric and epistemic uncertainty are fundamentally…

Machine Learning · Computer Science 2026-02-09 Sebastián Jiménez , Mira Jürgens , Willem Waegeman