English
Related papers

Related papers: A smoothing model for sample disclosure risk estim…

200 papers

We establish a simple connection between robust and differentially-private algorithms: private mechanisms which perform well with very high probability are automatically robust in the sense that they retain accuracy even if a constant…

Machine Learning · Statistics 2022-12-02 Kristian Georgiev , Samuel B. Hopkins

In recent years, numerous advances have been made in understanding how epidemic dynamics is affected by changes in individual behaviours. We propose an SIS-based compartmental model to tackle the simultaneous and coupled evolution of an…

Physics and Society · Physics 2025-11-04 Giulia de Meijere , Hugo Martin

Clinical prediction models enable healthcare professionals to estimate individual outcomes using patient characteristics. Current sample size guidelines for developing or updating models with continuous outcomes aim to minimise overfitting…

Although robust learning and local differential privacy are both widely studied fields of research, combining the two settings is just starting to be explored. We consider the problem of estimating a discrete distribution in total variation…

Statistics Theory · Mathematics 2022-04-21 Julien Chhor , Flore Sentenac

Population risk is always of primary interest in machine learning; however, learning algorithms only have access to the empirical risk. Even for applications with nonconvex nonsmooth losses (such as modern deep networks), the population…

Machine Learning · Computer Science 2018-10-19 Chi Jin , Lydia T. Liu , Rong Ge , Michael I. Jordan

Risk prediction models using genetic data have seen increasing traction in genomics. However, most of the polygenic risk models were developed using data from participants with similar (mostly European) ancestry. This can lead to biases in…

Machine Learning · Computer Science 2022-05-11 Prashnna K Gyawali , Yann Le Guen , Xiaoxia Liu , Hua Tang , James Zou , Zihuai He

We present a method for image-based crowd counting, one that can predict a crowd density map together with the uncertainty values pertaining to the predicted density map. To obtain prediction uncertainty, we model the crowd density values…

Computer Vision and Pattern Recognition · Computer Science 2020-10-06 Viresh Ranjan , Boyu Wang , Mubarak Shah , Minh Hoai

Quantitative studies in many fields involve the analysis of multivariate data of diverse types, including measurements that we may consider binary, ordinal and continuous. One approach to the analysis of such mixed data is to use a copula…

Statistics Theory · Mathematics 2007-06-13 Peter D. Hoff

Testing whether a sample survey is a credible representation of the population is an important question to ensure the validity of any downstream research. While this problem, in general, does not have an efficient solution, one might take a…

Machine Learning · Computer Science 2024-10-10 Debabrota Basu , Sourav Chakraborty , Debarshi Chanda , Buddha Dev Das , Arijit Ghosh , Arnab Ray

Network sampling is used around the world for surveys of vulnerable, hard-to-reach populations including people at risk for HIV, opioid misuse, and emerging epidemics. The sampling methods include tracing social links to add new people to…

Methodology · Statistics 2020-02-05 Steve Thompson

Graphs and networks are common ways of depicting biological information. In biology, many different biological processes are represented by graphs, such as regulatory networks, metabolic pathways and protein--protein interaction networks.…

Applications · Statistics 2010-11-16 Caiyan Li , Hongzhe Li

Population size estimates for hidden and hard-to-reach populations are particularly important when members are known to suffer from disproportion health issues or to pose health risks to the larger ambient population in which they are…

Social and Information Networks · Computer Science 2018-07-04 Bilal Khan , Hsuan-Wei Lee , Ian Fellows , Kirk Dombrowski

Research on formulae production in spreadsheets has established the practice as high risk yet unrecognised as such by industry. There are numerous software applications that are designed to audit formulae and find errors. However these are…

Human-Computer Interaction · Computer Science 2008-03-13 Simon Thorne , David Ball , Zoe Lawson

In applications where the study data are collected within cluster units (e.g., patients within transplant centers), it is often of interest to estimate and perform inference on the treatment effects of the cluster units. However, it is…

Applications · Statistics 2023-05-11 Nicholas Hartman , Kevin He

Cell lineage statistics is a powerful tool for inferring cellular parameters, such as division rate, death rate or the population growth rate. Yet, in practice such an analysis suffers from a basic problem: how should we treat incomplete…

Populations and Evolution · Quantitative Biology 2023-09-14 Arthur Genthon , Takashi Nozoe , Luca Peliti , David Lacoste

A statistical network model with overlapping communities can be generated as a superposition of mutually independent random graphs of varying size. The model is parameterized by the number of nodes, the number of communities, and the joint…

Probability · Mathematics 2024-12-19 Tommi Gröhn , Joona Karjalainen , Lasse Leskelä

In this paper, we consider the problem of partitioning a small data sample of size $n$ drawn from a mixture of $2$ sub-gaussian distributions. Our work is motivated by the application of clustering individuals according to their population…

Statistics Theory · Mathematics 2023-01-05 Shuheng Zhou

When collecting geocoded confidential data with the intent to disseminate, agencies often resort to altering the geographies prior to making data publicly available due to data privacy obligations. An alternative to releasing aggregated…

Methodology · Statistics 2019-05-14 Harrison Quick , Scott H. Holan , Christopher K. Wikle

Linear regression is a fundamental building block of statistical data analysis. It amounts to estimating the parameters of a linear model that maps input features to corresponding outputs. In the classical setting where the precision of…

Computer Science and Game Theory · Computer Science 2019-12-16 Nicolas Gast , Stratis Ioannidis , Patrick Loiseau , Benjamin Roussillon

Large samples have been generated routinely from various sources. Classic statistical models, such as smoothing spline ANOVA models, are not well equipped to analyze such large samples due to expensive computational costs. In particular,…

Methodology · Statistics 2020-04-23 Xiaoxiao Sun , Wenxuan Zhong , Ping Ma
‹ Prev 1 8 9 10 Next ›