中文
相关论文

相关论文: Robust Fusion Methods for Big Data

200 篇论文

We present a Bayesian data fusion method to approximate a posterior distribution from an ensemble of particle estimates that only have access to subsets of the data. Our approach relies on approximate probabilistic inference of model…

统计计算 · 统计学 2020-10-28 Caleb Miller , Michael D. Schneider , Jem N. Corcoran , Jason Bernstein

The era of huge data necessitates highly efficient machine learning algorithms. Many common machine learning algorithms, however, rely on computationally intensive subroutines that are prohibitively expensive on large datasets. Oftentimes,…

机器学习 · 计算机科学 2023-09-26 Mo Tiwari

In the analysis of large/big data sets, aggregation (replacing values of a variable over a group by a single value) is a standard way of reducing the size (complexity) of the data. Data analysis programs provide different aggregation…

机器学习 · 计算机科学 2023-03-29 Vladimir Batagelj

Most prognostic methods require a decent amount of data for model training. In reality, however, the amount of historical data owned by a single organization might be small or not large enough to train a reliable prognostic model. To…

机器学习 · 统计学 2024-04-11 Madi Arabi , Xiaolei Fang

Causal inference in modern largescale systems faces growing challenges, including highdimensional covariates, multi-valued treatments, massive observational (OBS) data, and limited randomized controlled trial (RCT) samples due to cost…

统计方法学 · 统计学 2026-02-27 Yuxi Du , Zhiheng Zhang , Haoxuan Li , Cong Fang , Jixing Xu , Peng Zhen , Jiecheng Guo

Fusion is a common tool for the analysis and utilization of available datasets and so an essential part of data mining and machine learning processes. However, a clear definition of the type of fusion is not always provided due to…

人工智能 · 计算机科学 2020-01-14 Silvia Beddar-Wiesing , Maarten Bieshaar

In a real expert system, one may have unreliable, unconfident, conflicting estimates of the value for a particular parameter. It is important for decision making that the information present in this aggregate somehow find its way into use.…

人工智能 · 计算机科学 2013-04-15 Henry Hamburger

This paper proposes a new theory and methodology to tackle the problem of unifying distributed analyses and inferences on shared parameters from multiple sources, into a single coherent inference. This surprisingly challenging problem…

统计方法学 · 统计学 2019-07-22 Hongsheng Dai , Murray Pollock , Gareth Roberts

Big Data are huge amounts of digital information that are automatically accrued or merged from several sources and rarely result from properly planned surveys. A Big Dataset is herein conceived of as a collection of information concerning a…

统计计算 · 统计学 2020-02-12 Deldossi Laura , Tommasi Chiara

The problems of computational data processing involving regression, interpolation, reconstruction and imputation for multidimensional big datasets are becoming more important these days, because of the availability of data and their widely…

统计方法学 · 统计学 2017-03-22 Yuri K. Shestopaloff , Alexander Y. Shestopaloff

Big data presents potential but unresolved value as a source for analysis and inference. However,selection bias, present in many of these datasets, needs to be accounted for so that appropriate inferences can be made on the target…

统计方法学 · 统计学 2025-01-09 Lyndon Ang , Robert Clark , Bronwyn Loong , Anders Holmberg

Big Data processing systems handle huge unstructured and structured data to store, process, and analyze through cluster analysis which helps in identifying unseen patterns to find the relationships between them. Clustering analysis over the…

分布式、并行与集群计算 · 计算机科学 2022-11-11 Dipesh Gyawali

The integration of real-world data (RWD) and randomized controlled trials (RCT) is increasingly important for advancing causal inference in scientific research. This combination holds great promise for enhancing the efficiency of causal…

统计方法学 · 统计学 2024-07-02 Xi Lin , Jens Magelund Tarp , Robin J. Evans

Deep learning requires regularization mechanisms to reduce overfitting and improve generalization. We address this problem by a new regularization method based on distributional robust optimization. The key idea is to modify the…

We develop large sample theory for merged data from multiple sources. Main statistical issues treated in this paper are (1) the same unit potentially appears in multiple datasets from overlapping data sources, (2) duplicated items are not…

统计理论 · 数学 2018-05-22 Takumi Saegusa

Data assimilation is the process to fuse information from priors, observations of nature, and numerical models, in order to obtain best estimates of the parameters or state of a physical system of interest. Presence of large errors in some…

数值分析 · 数学 2015-11-06 Vishwas Rao , Adrian Sandu , Michael Ng , Elias Nino-Ruiz

The inherent challenge of multimodal fusion is to precisely capture the cross-modal correlation and flexibly conduct cross-modal interaction. To fully release the value of each modality and mitigate the influence of low-quality multimodal…

机器学习 · 计算机科学 2023-06-07 Qingyang Zhang , Haitao Wu , Changqing Zhang , Qinghua Hu , Huazhu Fu , Joey Tianyi Zhou , Xi Peng

This paper proposes a new approach to multi-sensor data fusion. It suggests that aggregation of data from multiple sensors can be done more efficiently when we consider information about sensors' different characteristics. Similar to most…

系统与控制 · 电气工程与系统科学 2019-09-10 Mohammad Amin Ahmad Akhoundi , Ehsan Valavi

The bootstrap is a widely used procedure for statistical inference because of its simplicity and attractive statistical properties. However, the vanilla version of bootstrap is no longer feasible computationally for many modern massive…

统计方法学 · 统计学 2023-02-16 Yingying Ma , Chenlei Leng , Hansheng Wang

In recent years, multi-modal fusion has attracted a lot of research interest, both in academia, and in industry. Multimodal fusion entails the combination of information from a set of different types of sensors. Exploiting complementary…

机器学习 · 计算机科学 2020-08-27 Siddharth Roheda , Hamid Krim , Benjamin S. Riggan