English
Related papers

Related papers: Exploring pseudorandom value addition operations i…

200 papers

We introduce stochastic variational inference for Gaussian process models. This enables the application of Gaussian process (GP) models to data sets containing millions of data points. We show how GPs can be vari- ationally decomposed to…

Machine Learning · Computer Science 2013-09-27 James Hensman , Nicolo Fusi , Neil D. Lawrence

Multiscale transforms have become a key ingredient in many data processing tasks. With technological development, we observe a growing demand for methods to cope with non-linear data structures such as manifold values. In this paper, we…

Numerical Analysis · Mathematics 2021-08-17 Wael Mattar , Nir Sharon

Verifying that a statistically significant result is scientifically meaningful is not only good scientific practice, it is a natural way to control the Type I error rate. Here we introduce a novel extension of the p-value - a…

Methodology · Statistics 2018-07-04 Jeffrey D. Blume , Lucy DAgostino McGowan , William D. Dupont , Robert A. Greevy

Causal effect estimation from observational data is a crucial but challenging task. Currently, only a limited number of data-driven causal effect estimation methods are available. These methods either provide only a bound estimation of the…

Methodology · Statistics 2020-11-10 Debo Cheng , Jiuyong Li , Lin Liu , Kui Yu , Thuc Duy Lee , Jixue Liu

The study of the dynamics of the size of a population via mathematical modelling is a problem of interest and widely studied. Traditionally, continuous deterministic methods based on differential equations have been used to deal with this…

Probability · Mathematics 2020-01-08 J. -C. Cortés , A. Navarro-Quiles , J. -V. Romero , M. -D. Roselló

Machine Learning research, including work promoting fair or equitable algorithms, often relies on the concept of a data-generating probability distribution. The standard presumption is that since data points are 'sampled from' such a…

Machine Learning · Computer Science 2026-04-23 Benedikt Höltgen , Robert C. Williamson

Bayesian methods for learning Gaussian graphical models offer a principled framework for quantifying model uncertainty and incorporating prior knowledge. However, their scalability is constrained by the computational cost of jointly…

Methodology · Statistics 2025-08-28 Reza Mohammadi , Marit Schoonhoven , Lucas Vogels , S. Ilker Birbil

Data-driven anomaly detection methods typically build a model for the normal behavior of the target system, and score each data instance with respect to this model. A threshold is invariably needed to identify data instances with high (or…

Machine Learning · Statistics 2019-10-09 Sreelekha Guggilam , S. M. Arshad Zaidi , Varun Chandola , Abani Patra

Inspired by graph-based methodologies, we introduce a novel graph-spanning algorithm designed to identify changes in both offline and online data across low to high dimensions. This versatile approach is applicable to Euclidean and…

Machine Learning · Statistics 2026-01-09 Yang-Wen Sun , Katerina Papagiannouli , Vladimir Spokoiny

Recent work on the structure of social networks and the internet has focussed attention on graphs with distributions of vertex degree that are significantly different from the Poisson degree distributions that have been widely studied in…

Statistical Mechanics · Physics 2009-10-31 M. E. J. Newman , S. H. Strogatz , D. J. Watts

One important issue commonly encountered in the analysis of microarray data is to decide which and how many genes should be selected for further studies. For discriminant microarray data analyses based on statistical models, such as the…

Quantitative Methods · Quantitative Biology 2009-11-09 Wentian Li , Fengzhu Sun , Ivo Grosse

P-values are widely used in both the social and natural sciences to quantify the statistical significance of observed results. The recent surge of big data research has made the p-value an even more popular tool to test the significance of…

Applications · Statistics 2023-01-05 Bertie Vidgen , Taha Yasseri

We study how inherent randomness in the training process -- where each sample (or client in federated learning) contributes only to a randomly selected portion of training -- can be leveraged for privacy amplification. This includes (1)…

Machine Learning · Computer Science 2025-06-03 Andy Dong , Wei-Ning Chen , Ayfer Ozgur

An important question in statistical network analysis is how to estimate models of discrete and dependent network data with intractable likelihood functions, without sacrificing computational scalability and statistical guarantees. We…

Statistics Theory · Mathematics 2026-03-06 Jonathan R. Stewart , Michael Schweinberger

Cross-classified data frequently arise in scientific fields such as education, healthcare, and social sciences. A common modeling strategy is to introduce crossed random effects within a regression framework. However, this approach often…

Methodology · Statistics 2025-07-22 Shota Takeishi , Shonosuke Sugasawa

Differentially Private Synthetic Data Generation (DP-SDG) is a key enabler of private and secure tabular-data sharing, producing artificial data that carries through the underlying statistical properties of the input data. This typically…

Machine Learning · Computer Science 2025-04-16 Samuel Maddock , Shripad Gade , Graham Cormode , Will Bullock

Several recently developed methods have the potential to harness machine learning in the pursuit of target quantities inspired by causal inference, including inverse weighting, doubly robust estimating equations and substitution estimators…

Mixup is a widely adopted data augmentation technique known for enhancing the generalization of machine learning models by interpolating between data points. Despite its success and popularity, limited attention has been given to…

Machine Learning · Computer Science 2025-03-05 Chungpa Lee , Jongho Im , Joseph H. T. Kim

Finite mixture of Gaussian distributions provide a flexible semi-parametric methodology for density estimation when the variables under investigation have no boundaries. However, in practical applications variables may be partially bounded…

Methodology · Statistics 2019-12-30 Luca Scrucca

Semi-supervised learning (SSL) is a promising approach for training deep classification models using labeled and unlabeled datasets. However, existing SSL methods rely on a large unlabeled dataset, which may not always be available in many…

Machine Learning · Computer Science 2023-09-29 Shin'ya Yamaguchi