中文
相关论文

相关论文: Is External Information Useful for Data Fusion? An…

200 篇论文

Currently, the increase in financial returns from economic operations is constrained in view of the lack of a single efficiency criterion, which allows uniquely identify the business operation by their main feature - the possibility of…

最优化与控制 · 数学 2015-10-15 Igor Lutsenko

A fundamental aspect of statistics is the integration of data from different sources. Classically, Fisher and others were focused on how to integrate homogeneous (or only mildly heterogeneous) sets of data. More recently, as data are…

统计方法学 · 统计学 2024-03-20 Emily C. Hector , Ryan Martin

The information fusion field has recently been attracting a lot of interest within the scientific community, as it provides, through the combination of different sources of heterogeneous information, a fuller and/or more precise…

信息论 · 计算机科学 2025-10-28 Raúl Gutiérrez , Víctor Rampérez , Horacio Paggi , Juan A. Lara , Javier Soriano

Materials science data collection can be expensive, making the reuse and long-term utility of datasets critical important for future discovery campaigns. In practice, researchers prioritize a subset of properties due to research interests.…

Mutual Information (MI) is a powerful statistical measure that quantifies shared information between random variables, particularly valuable in high-dimensional data analysis across fields like genomics, natural language processing, and…

机器学习 · 计算机科学 2024-12-02 Andre O. Falcao

In this paper we focus on the estimation of mutual information from finite samples $(\mathcal{X}\times\mathcal{Y})$. The main concern with estimations of mutual information is their robustness under the class of transformations for which it…

数据分析、统计与概率 · 物理学 2020-02-04 Nicholas Carrara , Jesse Ernst

Randomized clinical trials are the gold standard for analyzing treatment effects, but high costs and ethical concerns can limit recruitment, potentially leading to invalid inferences. Incorporating external trial data with similar…

统计方法学 · 统计学 2024-09-09 Yujia Gu , Hanzhong Liu , Wei Ma

Transfer learning enhances prediction accuracy on a target distribution by leveraging data from a source distribution, demonstrating significant benefits in various applications. This paper introduces a novel dissimilarity measure that…

机器学习 · 统计学 2024-12-12 Mitsuhiro Fujikawa , Yohei Akimoto , Jun Sakuma , Kazuto Fukuchi

Data collection has become an increasingly important problem in robotic manipulation, yet there still lacks much understanding of how to effectively collect data to facilitate broad generalization. Recent works on large-scale robotic data…

机器人学 · 计算机科学 2024-05-22 Jensen Gao , Annie Xie , Ted Xiao , Chelsea Finn , Dorsa Sadigh

Recent advances in generative models facilitate the creation of synthetic data to be made available for research in privacy-sensitive contexts. However, the analysis of synthetic data raises a unique set of methodological challenges. In…

Mutual information (MI) is a fundamental measure of statistical dependence between two variables, yet accurate estimation from finite data remains notoriously difficult. No estimator is universally reliable, and common approaches fail in…

数据分析、统计与概率 · 物理学 2025-10-02 Eslam Abdelaleem , K. Michael Martini , Ilya Nemenman

Causal inference is central to many areas of artificial intelligence, including complex reasoning, planning, knowledge-base construction, robotics, explanation, and fairness. An active community of researchers develops and enhances…

人工智能 · 计算机科学 2019-11-05 Amanda Gentzel , Dan Garant , David Jensen

Synthetic data generation is a promising technique to facilitate the use of sensitive data while mitigating the risk of privacy breaches. However, for synthetic data to be useful in downstream analysis tasks, it needs to be of sufficient…

机器学习 · 统计学 2024-08-26 Thom Benjamin Volker , Peter-Paul de Wolf , Erik-Jan van Kesteren

Randomized controlled trials (RCTs) are the gold standard for evaluating causal effects but are often costly and difficult to scale; consequently, they are frequently augmented with auxiliary external controls in many applications. Prior…

统计方法学 · 统计学 2026-05-28 Jiawei Shan , Yiteng Tu , Guanbo Wang , Chao Ying , Jiwei Zhao

Data required to calibrate uncertain GCM parameterizations are often only available in limited regions or time periods, for example, observational data from field campaigns, or data generated in local high-resolution simulations. This…

应用统计 · 统计学 2022-10-05 Oliver R. A. Dunbar , Michael F. Howland , Tapio Schneider , Andrew M. Stuart

In the age of large and heterogeneous datasets, the integration of information from diverse sources is essential to improve parameter estimation. Multi-task learning offers a powerful approach by enabling simultaneous learning across…

统计方法学 · 统计学 2025-07-11 Sohom Bhattacharya , Yongzhuo Chen , Muxuan Liang

The maximum entropy principle can be used to assign utility values when only partial information is available about the decision maker's preferences. In order to obtain such utility values it is necessary to establish an analogy between…

统计金融 · 定量金融 2009-11-13 Andreia Dionisio , A. Heitor Reis

In this paper we propose an extension of the notion of deviation-based aggregation function tailored to aggregate multidimensional data. Our objective is both to improve the results obtained by other methods that try to select the best…

The era of big data has witnessed an increasing availability of multiple data sources for statistical analyses. We consider estimation of causal effects combining big main data with unmeasured confounders and smaller validation data with…

统计方法学 · 统计学 2021-08-24 Shu Yang , Peng Ding

We consider a data analyst's problem of purchasing data from strategic agents to compute an unbiased estimate of a statistic of interest. Agents incur private costs to reveal their data and the costs can be arbitrarily correlated with their…

计算机科学与博弈论 · 计算机科学 2018-09-06 Yiling Chen , Nicole Immorlica , Brendan Lucier , Vasilis Syrgkanis , Juba Ziani