中文
相关论文

相关论文: Discussion of "Data fission: splitting a single da…

200 篇论文

Our goal is to develop a general strategy to decompose a random variable $X$ into multiple independent random variables, without sacrificing any information about unknown parameters. A recent paper showed that for some well-known natural…

统计方法学 · 统计学 2025-12-23 Ameer Dharamshi , Anna Neufeld , Keshav Motwani , Lucy L. Gao , Daniela Witten , Jacob Bien

Classically, statistical datasets have a larger number of data points than features ($n > p$). The standard model of classical statistics caters for the case where data points are considered conditionally independent given the parameters.…

机器学习 · 统计学 2022-03-16 Sijia Li , Martín López-García , Neil D. Lawrence , Luisa Cutillo

Various phenomenological models of particle multiplicity distributions are discussed using a general form of the grand canonical partition function. These phenomenological models include a wide range of varied processes such as coherent…

核理论 · 物理学 2007-05-23 S. J. Lee , A. Z. Mekjian

Big data applications, such as medical imaging and genetics, typically generate datasets that consist of few observations n on many more variables p, a scenario that we denote as p>>n. Traditional data processing methods are often…

数据分析、统计与概率 · 物理学 2016-05-18 Magnus O. Ulfarsson , Frosti Palsson , Jakob Sigurdsson , Johannes R. Sveinsson

Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent…

分布式、并行与集群计算 · 计算机科学 2019-06-11 Salman Salloum , Yulin He , Joshua Zhexue Huang , Xiaoliang Zhang , Tamer Z. Emara , Chenghao Wei , Heping He

We aim to develop simultaneous inference tools for the mean function of functional data from sparse to dense. First, we derive a unified Gaussian approximation to construct simultaneous confidence bands of mean functions based on the…

统计方法学 · 统计学 2024-02-01 Leheng Cai , Qirui Hu

A common divide-and-conquer approach for Bayesian computation with big data is to partition the data, perform local inference for each piece separately, and combine the results to obtain a global posterior approximation. While being…

Standard Gaussian Process (GP) regression, a powerful machine learning tool, is computationally expensive when it is applied to large datasets, and potentially inaccurate when data points are sparsely distributed in a high-dimensional…

机器学习 · 计算机科学 2016-03-08 Z. Zhang , K. Duraisamy , N. A. Gumerov

In recent years, an increasing amount of data is collected in different and often, not cooperative, databases. The problem of privacy-preserving, distributed calculations over separated databases and, a relative to it, issue of private data…

数据库 · 计算机科学 2016-05-23 Philip Derbeko , Shlomi Dolev , Ehud Gudes , Jeffrey D. Ullman

This paper presents Sparse Partitioning, a Bayesian method for identifying predictors that either individually or in combination with others affect a response variable. The method is designed for regression problems involving binary or…

定量方法 · 定量生物学 2011-08-31 Doug Speed , Simon Tavaré

Advances in information technology have led to extremely large datasets that are often kept in different storage centers. Existing statistical methods must be adapted to overcome the resulting computational obstacles while retaining…

统计方法学 · 统计学 2021-11-12 Qiong Zhang , Jiahua Chen

Statistical inference on the mean of a Poisson distribution is a fundamentally important problem with modern applications in, e.g., particle physics. The discreteness of the Poisson distribution makes this problem surprisingly challenging,…

统计方法学 · 统计学 2012-07-03 Ryan Martin , Duncan Ermini Leaf , Chuanhai Liu

Semantic segmentation of aerial point cloud data can be utilised to differentiate which points belong to classes such as ground, buildings, or vegetation. Point clouds generated from aerial sensors mounted to drones or planes can utilise…

计算机视觉与模式识别 · 计算机科学 2022-11-30 Matthew Howe , Boris Repasky , Timothy Payne

A variety of machine learning tasks---e.g., matrix factorization, topic modelling, and feature allocation---can be viewed as learning the parameters of a probability distribution over bipartite graphs. Recently, a new class of models for…

机器学习 · 统计学 2017-12-07 Victor Veitch , Ekansh Sharma , Zacharie Naulet , Daniel M. Roy

Global data association is an essential prerequisite for robot operation in environments seen at different times or by different robots. Repetitive or symmetric data creates significant challenges for existing methods, which typically rely…

机器人学 · 计算机科学 2025-09-22 Yixuan Jia , Mason B. Peterson , Qingyuan Li , Yulun Tian , Jonathan P. How

We present a constructive and self-contained approach to data driven infinite partition-of-unity copulas that were recently introduced in the literature. In particular, we consider negative binomial and Poisson copulas and present a…

风险管理 · 定量金融 2020-12-17 Dietmar Pfeifer , Andreas Mändle , Olena Ragulina

A new generalization of the family of Poisson-G is called beta Poisson-G family of distribution. Useful expansions of the probability density function and the cumulative distribution function of the proposed family are derived and seen as…

统计理论 · 数学 2020-05-22 Laba Handique , Subrata Chakraborty , Farrukh Jamal

A representation of Gaussian distributed sparsely sampled longitudinal data in terms of predictive distributions for their functional principal component scores (FPCs) maps available data for each subject to a multivariate Gaussian…

统计方法学 · 统计学 2026-03-13 Álvaro Gajardo , Xiongtao Dai , Hans-Georg Müller

Modern technologies are generating ever-increasing amounts of data. Making use of these data requires methods that are both statistically sound and computationally efficient. Typically, the statistical and computational aspects are treated…

统计方法学 · 统计学 2022-09-15 Mahsa Taheri , Néhémy Lim , Johannes Lederer

In this work, we develop a method named Twinning, for partitioning a dataset into statistically similar twin sets. Twinning is based on SPlit, a recently proposed model-independent method for optimally splitting a dataset into training and…

机器学习 · 统计学 2022-02-17 Akhil Vakayil , V. Roshan Joseph