English
Related papers

Related papers: bayesgrid: An Open-Source Python Tool for Generati…

200 papers

The generation of synthetic data is an essential tool to study complex systems, allowing for example to test models of these in precisely controlled settings, or to parametrize simulation models when data is missing. This paper focuses on…

Applications · Statistics 2019-11-25 Juste Raimbault

This paper introduces PyProD, a Python-based machine learning (ML)-compatible test-bed for evaluating the efficacy of protection schemes in electric distribution grids. This testbed is designed to bridge the gap between conventional power…

Systems and Control · Electrical Eng. & Systems 2021-09-14 Dongqi Wu , Dileep Kalathil , Miroslav Begovic , Le Xie

We study the problem of probability distribution matching and sampling on near-term quantum computers, aiming to construct parameterized circuits that generate samples from a target distribution while minimizing resource overhead. This task…

Quantum Physics · Physics 2026-05-26 Nicholas S. DiBrita , Jason Han , Krishna Bhatia , Younghyun Cho , Hengrui Luo , Tirthak Patel

Synthetic datasets are important for evaluating and testing machine learning models. When evaluating real-life recommender systems, high-dimensional categorical (and sparse) datasets are often considered. Unfortunately, there are not many…

Information Retrieval · Computer Science 2024-12-11 Miha Malenšek , Blaž Škrlj , Blaž Mramor , Jure Demšar

Replicated network data are increasingly available in many research fields. In connectomic applications, inter-connections among brain regions are collected for each patient under study, motivating statistical models which can flexibly…

Methodology · Statistics 2018-09-11 Daniele Durante , David B. Dunson , Joshua T. Vogelstein

We present a Bayesian model for estimating the joint distribution of multivariate categorical data when units are nested within groups. Such data arise frequently in social science settings, for example, people living in households. The…

Methodology · Statistics 2016-10-31 Jingchen Hu , Jerome P. Reiter , Quanli Wang

Datasets with hundreds of variables and many missing values are commonplace. In this setting, it is both statistically and computationally challenging to detect true predictive relationships between variables and also to suppress false…

Machine Learning · Statistics 2018-04-03 Feras Saad , Vikash Mansinghka

This paper motivates the use of random-bridges -- stochastic processes conditioned to take target distributions at fixed timepoints -- in the realm of generative modelling. Herein, random-bridges can act as stochastic transports between two…

Machine Learning · Computer Science 2026-04-07 Stefano Goria , Levent A. Mengütürk , Murat C. Mengütürk , Berkan Sesen

The evolution of communities in dynamic (time-varying) network data is a prominent topic of interest. A popular approach to understanding these dynamic networks is to embed the dyadic relations into a latent metric space. While methods for…

Methodology · Statistics 2020-03-18 Joshua Daniel Loyal , Yuguo Chen

I present StarEstate, an open-source Python package for producing rapid, statistically robust galactic population synthesis models. By utilizing optimized pre-calculated inverse-cumulative distribution function samplers, the tool generates…

Instrumentation and Methods for Astrophysics · Physics 2025-11-27 Amedeo Romagnolo

Quantifying spatial and/or temporal associations in multivariate geolocated data of different types is achievable via spatial random effects in a Bayesian hierarchical model, but severe computational bottlenecks arise when spatial…

Methodology · Statistics 2024-04-02 Michele Peruzzi , David B. Dunson

In this paper, we propose a novel method for generating a synthetic dataset obeying Gaussian distribution. Compared to the commonly used benchmark datasets with unknown distribution, the synthetic dataset has an explicit distribution, i.e.,…

Computer Vision and Pattern Recognition · Computer Science 2019-07-01 Xinjie Lan

Evaluating time series attribution methods is difficult because real-world datasets rarely provide ground truth for which time points drive a prediction. A common workaround is to generate synthetic data where class-discriminating features…

Machine Learning · Computer Science 2026-03-10 Gregor Baer

The increasing complexity of the power grid, due to higher penetration of distributed resources and the growing availability of interconnected, distributed metering devices re- quires novel tools for providing a unified and consistent view…

Machine Learning · Statistics 2017-05-25 Francesco Fusco , Seshu Tirupathi , Robert Gormally

Recent years have noticed an increasing interest among academia and industry towards analyzing the electrical consumption of residential buildings and employing smart home energy management systems (HEMS) to reduce household energy…

Machine Learning · Computer Science 2023-05-17 Mina Razghandi , Hao Zhou , Melike Erol-Kantarci , Damla Turgut

We present a distributionally robust PAC-Bayesian framework for certifying the performance of learning-based finite-horizon controllers. While existing PAC-Bayes control literature typically assumes bounded losses and matching training and…

Machine Learning · Computer Science 2026-04-14 Domagoj Herceg , Duarte Antunes

We propose the Gaussian-Linear Hidden Markov model (GLHMM), a generalisation of different types of HMMs commonly used in neuroscience. In short, the GLHMM is a general framework where linear regression is used to flexibly parameterise the…

Neurons and Cognition · Quantitative Biology 2024-10-02 Diego Vidaurre , Laura Masaracchia , Nick Y. Larsen , Lenno R. P. T Ruijters , Sonsoles Alonso , Christine Ahrends , Mark W. Woolrich

We propose two methods for generating non-Gaussian maps with fixed power spectrum and bispectrum. The first makes use of a recently proposed rigorous, non-perturbative, Bayesian framework for generating non-Gaussian distributions. The…

Astrophysics · Physics 2009-11-06 Carlo R. Contaldi , Joao Magueijo

Data is the fuel of data science and machine learning techniques for smart grid applications, similar to many other fields. However, the availability of data can be an issue due to privacy concerns, data size, data quality, and so on. To…

Machine Learning · Computer Science 2022-01-20 Mina Razghandi , Hao Zhou , Melike Erol-Kantarci , Damla Turgut

Generating graph-structured data is crucial in applications such as molecular generation, knowledge graphs, and network analysis. However, their discrete, unordered nature makes them difficult for traditional generative models, leading to…

Machine Learning · Computer Science 2026-04-14 Ole Petersen , Marcel Kollovieh , Marten Lienen , Stephan Günnemann