English
Related papers

Related papers: Data fission: splitting a single data point

200 papers

Approximation algorithms are widely used in many engineering problems. To obtain a data set for approximation a factorial design of experiments is often used. In such case the size of the data set can be very large. Therefore, one of the…

Methodology · Statistics 2014-07-04 Mikhail Belyaev , Evgeny Burnaev , Yermek Kapushev

As a class of generative artificial intelligence frameworks inspired by statistical physics, diffusion models have shown extraordinary performance in synthesizing complicated data distributions through a denoising process gradually guided…

Machine Learning · Computer Science 2026-04-23 Fangjun Hu , Guangkuo Liu , Yifan F. Zhang , Xun Gao

This work provides a new multinomial resampling procedure for particle filter resampling, focused on the case where the number of samples required is less than or equal to the size of the underlying discrete distribution. This setting is…

Data Structures and Algorithms · Computer Science 2026-04-03 Andrey A. Popov

This paper presents a new approach for Gaussian process (GP) regression for large datasets. The approach involves partitioning the regression input domain into multiple local regions with a different local GP model fitted in each region.…

Machine Learning · Computer Science 2018-07-10 Chiwoo Park , Daniel Apley

First Few X (FFX) studies collect household-stratified data in the early stages of a pandemic, in order to infer severity and transmissibility of an emerging disease. We present a Bayesian method to approximately infer population-level…

Populations and Evolution · Quantitative Biology 2021-11-22 P. G. Ballard , A. J. Black , J. V. Ross

Deep Gaussian Processes learn probabilistic data representations for supervised learning by cascading multiple Gaussian Processes. While this model family promises flexible predictive distributions, exact inference is not tractable.…

Machine Learning · Statistics 2020-10-23 Jakob Lindinger , David Reeb , Christoph Lippert , Barbara Rakitsch

No polynomial-time algorithm is known to test whether a sparse polynomial G divides another sparse polynomial $F$. While computing the quotient Q=F quo G can be done in polynomial time with respect to the sparsities of F, G and Q, this is…

Symbolic Computation · Computer Science 2021-07-21 Pascal Giorgi , Bruno Grenet , Armelle Perret du Cray

We consider testing whether a set of Gaussian variables, selected from the data, is independent of the remaining variables. We assume that this set is selected via a very simple approach that is commonly used across scientific disciplines:…

Methodology · Statistics 2022-11-04 Arkajyoti Saha , Daniela Witten , Jacob Bien

In this work, we develop a method named Twinning, for partitioning a dataset into statistically similar twin sets. Twinning is based on SPlit, a recently proposed model-independent method for optimally splitting a dataset into training and…

Machine Learning · Statistics 2022-02-17 Akhil Vakayil , V. Roshan Joseph

In this work we consider time series with a finite number of discrete point changes. We assume that the data in each segment follows a different probability density functions (pdf). We focus on the case where the data in all segments are…

Data Analysis, Statistics and Probability · Physics 2007-05-23 Ali Mohammad-Djafari , Olivier Feron

The particle-in-cell numerical method of plasma physics balances a trade-off between computational cost and intrinsic noise. Inference on data produced by these simulations generally consists of binning the data to recover the particle…

Plasma Physics · Physics 2022-02-03 John Donaghy , Kai Germaschewski

Many statistical estimands of interest (e.g., in regression or causality) are functions of the joint distribution of multiple random variables. But in some applications, data is not available that measures all random variables on each…

Methodology · Statistics 2025-02-11 Yicong Jiang , Lucas Janson

The Dantzig selector is a widely used and effective method for variable selection in ultra-high-dimensional data. Feature splitting is an efficient processing technique that involves dividing these ultra-high-dimensional variable datasets…

Computation · Statistics 2025-04-04 Xiaofei Wu , Yue Chao , Rongmei Liang , Shi Tang , Zhiming Zhang

Tangles were originally introduced as a concept to formalize regions of high connectivity in graphs. In recent years, they have also been discovered as a link between structural graph theory and data science: when interpreting similarity in…

Statistics Theory · Mathematics 2024-03-12 Eva Fluck , Sandra Kiefer , Christoph Standke

Computing accurate estimates of the Fourier transform of analog signals from discrete data points is important in many fields of science and engineering. The conventional approach of performing the discrete Fourier transform of the data…

Machine Learning · Statistics 2017-12-08 Luca Ambrogioni , Eric Maris

Nonuniform subsampling methods are effective to reduce computational burden and maintain estimation efficiency for massive data. Existing methods mostly focus on subsampling with replacement due to its high computational efficiency. If the…

Methodology · Statistics 2021-07-06 Jun Yu , HaiYing Wang , Mingyao Ai , Huiming Zhang

In this work we tackle the problem of estimating the density $f_X$ of a random variable $X$ by successive smoothing, such that the smoothed random variable $Y$ fulfills $(\partial_t - \Delta_1)f_Y(\,\cdot\,, t) = 0$, $f_Y(\,\cdot\,, 0) =…

Computer Vision and Pattern Recognition · Computer Science 2023-10-20 Martin Zach , Thomas Pock , Erich Kobler , Antonin Chambolle

Post-selection inference (PoSI) is a statistical technique for obtaining valid confidence intervals and p-values when hypothesis generation and testing use the same source of data. PoSI can be used on a range of popular algorithms including…

Methodology · Statistics 2023-05-23 Erik Drysdale

Post-data statistical inference concerns making probability statements about model parameters conditional on observed data. When a priori knowledge about parameters is available, post-data inference can be conveniently made from Bayesian…

Statistics Theory · Mathematics 2025-06-05 Yang Liu , Jan Hannig , Alexander C Murph

Motivated by the need for computationally tractable spatial methods in neuroimaging studies, we develop a distributed and integrated framework for estimation and inference of Gaussian process model parameters with ultra-high-dimensional…

Methodology · Statistics 2024-02-09 Emily C. Hector , Brian J. Reich , Ani Eloyan