English
Related papers

Related papers: A smoothing approach for masking spatial data

200 papers

In this paper, our goal is to generate synthetic data for heterogeneous (mixed-type) tabular datasets with high machine learning utility (MLu). Since the MLu performance depends on accurately approximating the conditional distributions, we…

Machine Learning · Computer Science 2024-08-20 Seunghwan An , Gyeongdong Woo , Jaesung Lim , ChangHyun Kim , Sungchul Hong , Jong-June Jeon

The characteristics and determinants of health and disease are often organised in space, reflecting our spatially extended nature. Understanding the influence of such factors requires models capable of capturing spatial relations. Though a…

The growing volume of data usually creates an interesting challenge for the need of data analysis tools that discover regularities in these data. Data mining has emerged as disciplines that contribute tools for data analysis, discovery of…

Databases · Computer Science 2011-08-30 Abhishek Taneja , R. K. Chauhan

We consider the problem of estimating a spatially varying density function, motivated by problems that arise in large-scale radiological survey and anomaly detection. In this context, the density functions to be estimated are the background…

Methodology · Statistics 2017-11-17 Wesley Tansey , Alex Athey , Alex Reinhart , James G. Scott

In this paper we set out general principles and develop geostatistical methods for the analysis of data from spatio-temporally referenced prevalence surveys. Our objective is to provide a tutorial guide that can be used in order to identify…

Methodology · Statistics 2018-02-20 Emanuele Giorgi , Peter J. Diggle , Robert W. Snow , Abdisalan M. Noor

Regularization is an effective way to promote the generalization performance of machine learning models. In this paper, we focus on label smoothing, a form of output distribution regularization that prevents overfitting of a neural network…

Machine Learning · Computer Science 2020-07-07 Weizhi Li , Gautam Dasarathy , Visar Berisha

Masked (or absorbing) diffusion is actively explored as an alternative to autoregressive models for generative modeling of discrete data. However, existing work in this area has been hindered by unnecessarily complex model formulations and…

Machine Learning · Computer Science 2025-01-17 Jiaxin Shi , Kehang Han , Zhe Wang , Arnaud Doucet , Michalis K. Titsias

The ubiquity of missing data in urban intelligence systems, attributable to adverse environmental conditions and equipment failures, poses a significant challenge to the efficacy of downstream applications, notably in the realms of traffic…

Machine Learning · Computer Science 2026-05-25 Songyu Ke , Chenyu Wu , Yuxuan Liang , Huiling Qin , Junbo Zhang , Yu Zheng

In almost any geostatistical analysis, one of the underlying, often implicit, modelling assump- tions is that the spatial locations, where measurements are taken, are recorded without error. In this study we develop geostatistical inference…

Applications · Statistics 2017-11-02 Claudio Fronterrè , Emanuele Giorgi , Peter J. Diggle

Generative machine learning models are being increasingly viewed as a way to share sensitive data between institutions. While there has been work on developing differentially private generative modeling approaches, these approaches…

Cryptography and Security · Computer Science 2022-10-13 Yixi Xu , Sumit Mukherjee , Xiyang Liu , Shruti Tople , Rahul Dodhia , Juan Lavista Ferres

The application of deep neural networks in geospatial data has become a trending research problem in the present day. A significant amount of statistical research has already been introduced, such as generalized least square optimization by…

Machine Learning · Statistics 2024-11-07 Debjoy Thakur

Data sharing ecosystems connect providers, consumers, and intermediaries to facilitate the exchange and use of data for a wide range of downstream tasks. In sensitive domains such as healthcare, privacy is enforced as a hard constraint, any…

Motivated by two case studies using primary care records from the Clinical Practice Research Datalink, we describe statistical methods that facilitate the analysis of tall data, with very large numbers of observations. Our focus is on…

Methodology · Statistics 2018-05-14 Kirsty Rhodes , Rebecca Turner , Rupert Payne , Ian White

Disease maps display the spatial pattern in disease risk, so that high-risk clusters can be identified. The spatial structure in the risk map is typically represented by a set of random effects, which are modelled with a conditional…

Methodology · Statistics 2015-03-20 Duncan Lee

In today's modern era of Big data, computationally efficient and scalable methods are needed to support timely insights and informed decision making. One such method is sub-sampling, where a subset of the Big data is analysed and used as…

Methodology · Statistics 2022-09-07 Amalan Mahendran , Helen Thompson , James M. McGree

Prevalence mapping in low resource settings is an increasingly important endeavor to guide policy making and to spatially and temporally characterize the burden of disease. We will focus our discussion on consideration of the complex design…

Methodology · Statistics 2016-08-15 Jon Wakefield , Daniel Simpson , Jessica Godwin

Spatial data is a rich source of information for actuarial applications: knowledge of a risk's location could improve an insurance company's ratemaking, reserving or risk management processes. Insurance companies with high exposures in a…

Applications · Statistics 2024-08-22 Christopher Blier-Wong , Hélène Cossette , Luc Lamontagne , Etienne Marceau

As one of the central tasks in machine learning, regression finds lots of applications in different fields. An existing common practice for solving regression problems is the mean square error (MSE) minimization approach or its regularized…

Machine Learning · Statistics 2022-11-24 Jirong Yi , Qiaosheng Zhang , Zhen Chen , Qiao Liu , Wei Shao , Yusen He , Yaohua Wang

There is a growing need for flexible general frameworks that integrate individual-level data with external summary information for improved statistical inference. External information relevant for a risk prediction model may come in…

Methodology · Statistics 2023-04-11 Tian Gu , Jeremy M. G. Taylor , Bhramar Mukherjee

Hazard models are the most commonly used tool to analyse time-to-event data. If more than one time scale is relevant for the event under study, models are required that can incorporate the dependence of a hazard along two (or more) time…

Methodology · Statistics 2025-01-14 Angela Carollo , Paul H. C. Eilers , Hein Putter , Jutta Gampe