English
Related papers

Related papers: Copula-based transferable models for synthetic pop…

200 papers

Copula-based models provide a great deal of flexibility in modelling multivariate distributions, allowing for the specifications of models for the marginal distributions separately from the dependence structure (copula) that links them to…

Methodology · Statistics 2021-09-09 Nicolás Kuschinski , Alejandro Jara

Copula modelling has become ubiquitous in modern statistics. Here, the problem of nonparametrically estimating a copula density is addressed. Arguably the most popular nonparametric density estimator, the kernel estimator is not suitable…

Methodology · Statistics 2014-04-18 Gery Geenens , Arthur Charpentier , Davy Paindaveine

Generative Policy-based Models aim to enable a coalition of systems, be they devices or services to adapt according to contextual changes such as environmental factors, user preferences and different tasks whilst adhering to various…

Artificial Intelligence · Computer Science 2019-05-01 Daniel Cunnington , Graham White , Geeth de Mel

Motivated by challenges in the analysis of biomedical data and observational studies, we develop statistical boosting for the general class of bivariate distributional copula regression with arbitrary marginal distributions, which is suited…

Methodology · Statistics 2024-03-05 Guillermo Briseño Sanchez , Nadja Klein , Hannah Klinkhammer , Andreas Mayr

In this work, we propose a non-iterative Gaussian transformation strategy based on copula function, which doesn't require some commonly seen restrictive assumptions in the previous studies such as the elliptically symmetric distribution…

Methodology · Statistics 2022-03-29 Rongxiang Rui , Maozai Tian

Dependence strucuture estimation is one of the important problems in machine learning domain and has many applications in different scientific areas. In this paper, a theoretical framework for such estimation based on copula and copula…

Machine Learning · Computer Science 2019-09-11 Jian Ma , Zengqi Sun

The generation of synthetic data is an essential tool to study complex systems, allowing for example to test models of these in precisely controlled settings, or to parametrize simulation models when data is missing. This paper focuses on…

Applications · Statistics 2019-11-25 Juste Raimbault

Electronic health records (EHR) often contain different rates of representation of certain subpopulations (SP). Factors like patient demographics, clinical condition prevalence, and medical center type contribute to this…

Machine Learning · Computer Science 2024-03-12 Oriel Perets , Nadav Rappoport

Local climate information is crucial for impact assessment and decision-making, yet coarse global climate simulations cannot capture small-scale phenomena. Current statistical downscaling methods infer these phenomena as temporally…

Machine Learning · Computer Science 2025-09-24 Jonathan Schmidt , Luca Schmidt , Felix Strnad , Nicole Ludwig , Philipp Hennig

Instance-wise feature selection and ranking methods can achieve a good selection of task-friendly features for each sample in the context of neural networks. However, existing approaches that assume feature subsets to be independent are…

Machine Learning · Computer Science 2023-08-02 Hanyu Peng , Guanhua Fang , Ping Li

We are studying the problems of modeling and inference for multivariate count time series data with Poisson marginals. The focus is on linear and log-linear models. For studying the properties of such processes we develop a novel conceptual…

Methodology · Statistics 2017-04-10 Paul Doukhan , Konstantinos Fokianos , Bård Støve , Dag Tjøstheim

This paper addresses the problem of quantification and propagation of uncertainties associated with dependence modeling when data for characterizing probability models are limited. Practically, the system inputs are often assumed to be…

Computation · Statistics 2020-04-14 Jiaxin Zhang , Michael D. Shields

Correctly capturing the symmetry transformations of data can lead to efficient models with strong generalization capabilities, though methods incorporating symmetries often require prior knowledge. While recent advancements have been made…

Data augmentation via synthetic data generation has been shown to be effective in improving model performance and robustness in the context of scarce or low-quality data. Using the data valuation framework to statistically identify…

Machine Learning · Computer Science 2025-02-11 Tommaso Ferracci , Leonie Tabea Goldmann , Anton Hinel , Francesco Sanna Passino

Conditional generative models map input variables to complex, high-dimensional distributions, enabling realistic sample generation in a diverse set of domains. A critical challenge with these models is the absence of calibrated uncertainty,…

Machine Learning · Computer Science 2026-02-02 Qidong Yang , Qianyu Julie Zhu , Jonathan Giezendanner , Youssef Marzouk , Stephen Bates , Sherrie Wang

We develop adaptive estimation and inference methods for high-dimensional Gaussian copula regression that achieve the same performance without the knowledge of the marginal transformations as that for high-dimensional linear regression.…

Methodology · Statistics 2015-12-09 T. Tony Cai , Linjun Zhang

Generative models inspired by dynamical transport of measure -- such as flows and diffusions -- construct a continuous-time map between two probability densities. Conventionally, one of these is the target density, only accessible through…

Machine Learning · Computer Science 2024-09-24 Michael S. Albergo , Mark Goldstein , Nicholas M. Boffi , Rajesh Ranganath , Eric Vanden-Eijnden

What is a population? This review considers how a population may be defined in terms of understanding the structure of the underlying genetics of the individuals involved. The main approach is to consider statistically identifiable groups…

Populations and Evolution · Quantitative Biology 2013-06-05 Daniel John Lawson

We describe here a new method to estimate copula measure. From N observations of two variables X and Y, we draw a huge number m of subsamples (size n<N), and we compute the joint ranks in these subsamples. Then, for each bivariate rank…

Methodology · Statistics 2007-09-26 Jérôme Collet

Quantification of microbial interactions from 16S rRNA and meta-genomic sequencing data is difficult due to their sparse nature, as well as the fact that the data only provides measures of relative abundance. In this paper, we propose using…

Methodology · Statistics 2021-11-04 Rebecca A. Deek , Hongzhe Li