English
Related papers

Related papers: Is External Information Useful for Data Fusion? An…

200 papers

This paper provides a systematic approach to semiparametric identification that is based on statistical information as a measure of its "quality". Identification can be regular or irregular, depending on whether the Fisher information for…

Statistics Theory · Mathematics 2021-07-01 Juan Carlos Escanciano

We develop a data-driven information-theoretic framework for sharp partial identification of causal effects under unmeasured confounding. Existing approaches often rely on restrictive assumptions, such as bounded or discrete outcomes;…

Machine Learning · Statistics 2026-02-24 Yonghan Jung , Bogyeong Kang

We propose a semiparametric data fusion framework for efficient inference on survival probabilities by integrating right-censored and current status data. Existing data fusion methods focus largely on fusing right-censored data only, while…

Methodology · Statistics 2025-09-15 Xiudi Li , Sijia Li

Entropy is a measure of heterogeneity widely used in applied sciences, often when data are collected over space. Recently, a number of approaches has been proposed to include spatial information in entropy. The aim of entropy is to…

Statistics Theory · Mathematics 2019-11-12 Linda Altieri , Daniela Cocchi , Giulia Roli

In the era of big data, the explosive growth of multi-source heterogeneous data offers many exciting challenges and opportunities for improving the inference of conditional average treatment effects. In this paper, we investigate…

Machine Learning · Statistics 2022-11-02 Xinyu Li , Yilin Li , Qing Cui , Longfei Li , Jun Zhou

Effective summarisation evaluation metrics enable researchers and practitioners to compare different summarisation systems efficiently. Estimating the effectiveness of an automatic evaluation metric, termed meta-evaluation, is a critically…

Computation and Language · Computer Science 2024-10-01 Xiang Dai , Sarvnaz Karimi , Biaoyan Fang

Measuring model performance is a key issue for deep learning practitioners. However, we often lack the ability to explain why a specific architecture attains superior predictive accuracy for a given data set. Often, validation accuracy is…

Machine Learning · Computer Science 2022-04-01 Marius C. Landverk , Signe Riemer-Sørensen

Generalist imitation learning policies trained on large datasets show great promise for solving diverse manipulation tasks. However, to ensure generalization to different conditions, policies need to be trained with data collected across a…

The concepts of information transfer and causal effect have received much recent attention, yet often the two are not appropriately distinguished and certain measures have been suggested to be suitable for both. We discuss two existing…

Adaptation and Self-Organizing Systems · Physics 2012-03-05 Joseph T. Lizier , Mikhail Prokopenko

Long memory in the sense of slowly decaying autocorrelations is a stylized fact in many time series from economics and finance. The fractionally integrated process is the workhorse model for the analysis of these time series. Nevertheless,…

Econometrics · Economics 2023-09-22 Uwe Hassler , Marc-Oliver Pohle

Modern data is messy and high-dimensional, and it is often not clear a priori what are the right questions to ask. Instead, the analyst typically needs to use the data to search for interesting analyses to perform and hypotheses to test.…

Machine Learning · Statistics 2019-10-09 Daniel Russo , James Zou

Integrating probability and non-probability samples is increasingly important, yet unknown sampling mechanisms in non-probability sources complicate identification and efficient estimation. We develop semiparametric theory for dual-frame…

Methodology · Statistics 2026-01-14 Kosuke Morikawa , Jae Kwang Kim

The increasing availability of passively observed data has yielded a growing methodological interest in "data fusion." These methods involve merging data from observational and experimental sources to draw causal conclusions -- and they…

Methodology · Statistics 2021-12-15 Evan Rosenman , Art B. Owen

Time-to-event analyses are often plagued by both -- possibly unmeasured -- confounding and competing risks. To deal with the former, the use of instrumental variables for effect estimation is rapidly gaining ground. We show how to make use…

Methodology · Statistics 2018-01-04 Torben Martinussen , Stijn Vansteelandt

We consider the task of identifying and estimating a parameter of interest in settings where data is missing not at random (MNAR). In general, such parameters are not identified without strong assumptions on the missing data model. In this…

Methodology · Statistics 2024-02-29 Zixiao Wang , AmirEmad Ghassami , Ilya Shpitser

Estimating causal effects of joint interventions on multiple variables is crucial in many domains, but obtaining data from such simultaneous interventions can be challenging. Our study explores how to learn joint interventional effects…

Machine Learning · Statistics 2025-06-06 Armin Kekić , Sergio Hernan Garrido Mejia , Bernhard Schölkopf

We study the problem of multi-task non-smooth optimization that arises ubiquitously in statistical learning, decision-making and risk management. We develop a data fusion approach that adaptively leverages commonalities among a large number…

Machine Learning · Statistics 2022-10-25 Henry Lam , Kaizheng Wang , Yuhang Wu , Yichen Zhang

Record fusion is the task of aggregating multiple records that correspond to the same real-world entity in a database. We can view record fusion as a machine learning problem where the goal is to predict the "correct" value for each…

Machine Learning · Computer Science 2020-06-19 Alireza Heidari , George Michalopoulos , Shrinu Kushagra , Ihab F. Ilyas , Theodoros Rekatsinas

Mutual information is widely used, in a descriptive way, to measure the stochastic dependence of categorical random variables. In order to address questions such as the reliability of the descriptive value, one must consider…

Machine Learning · Computer Science 2007-07-13 Marcus Hutter , Marco Zaffalon

Semivalue-based data valuation uses cooperative-game theory intuitions to assign each data point a value reflecting its contribution to a downstream task. Still, those values depend on the practitioner's choice of utility, raising the…

Artificial Intelligence · Computer Science 2026-03-11 Mélissa Tamine , Benjamin Heymann , Maxime Vono , Patrick Loiseau
‹ Prev 1 8 9 10 Next ›