English
Related papers

Related papers: Hidden population size estimation from respondent-…

200 papers

Estimation of population size using incomplete lists (also called the capture-recapture problem) has a long history across many biological and social sciences. For example, human rights and other groups often construct partial and…

Methodology · Statistics 2021-08-03 Manjari Das , Edward H. Kennedy , Nicholas P. Jewell

An innovative sampling strategy is proposed, which applies to large-scale population-based surveys targeting a rare trait that is unevenly spread over a geographical area of interest. Our proposal is characterised by the ability to tailor…

Methodology · Statistics 2020-04-07 Fulvia Mecatti , Charalambos Sismanidis , Emanuela Furfaro

Respondent-Driven Sampling is a popular technique for sampling hidden populations. This paper models Respondent-Driven Sampling as a Markov process indexed by a tree. Our main results show that the Volz-Heckathorn estimator is…

Methodology · Statistics 2016-08-30 Xiao Li , Karl Rohe

Social media provide access to behavioural data at an unprecedented scale and granularity. However, using these data to understand phenomena in a broader population is difficult due to their non-representativeness and the bias of…

Computers and Society · Computer Science 2019-05-16 Zijian Wang , Scott A. Hale , David Adelani , Przemyslaw A. Grabowicz , Timo Hartmann , Fabian Flöck , David Jurgens

In this study, we propose a method Distributionally Robust Safe Screening (DRSS), for identifying unnecessary samples and features within a DR covariate shift setting. This method effectively combines DR learning, a paradigm aimed at…

Nonprobability samples have rapidly emerged to address time-sensitive priority topics in a variety of fields. While these data are timely, they are prone to selection bias. To mitigate selection bias, a large number of survey research…

Methodology · Statistics 2025-08-08 Kangrui Liu , Lingxiao Wang , Yan Li

Randomized algorithms are used in many state-of-the-art solvers for constraint satisfaction problems (CSP) and Boolean satisfiability (SAT) problems. For many of these problems, there is no single solver which will dominate others. Having…

Machine Learning · Computer Science 2021-06-25 Jake Tuero , Michael Buro

Methods for carefully selecting or generating a small set of training data to learn from, i.e., data pruning, coreset selection, and data distillation, have been shown to be effective in reducing the ever-increasing cost of training neural…

We consider the problem of privacy protection in Reinforcement Learning (RL) algorithms that operate over population processes, a practical but understudied setting that includes, for example, the control of epidemics in large populations…

Machine Learning · Computer Science 2024-06-26 Samuel Yang-Zhao , Kee Siong Ng

Accurately analyzing graph properties of social networks is a challenging task because of access limitations to the graph data. To address this challenge, several algorithms to obtain unbiased estimates of properties from few samples via a…

Social and Information Networks · Computer Science 2020-07-14 Kazuki Nakajima , Kazuyuki Shudo

Detection-based methods have been viewed unfavorably in crowd analysis due to their poor performance in dense crowds. However, we argue that the potential of these methods has been underestimated, as they offer crucial information for crowd…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Shaokai Wu , Fengyu Yang

Randomized Response (RR) is a protocol designed to collect and analyze categorical data with local differential privacy guarantees. It has been used as a building block of mechanisms deployed by Big tech companies to collect app or web…

Cryptography and Security · Computer Science 2026-01-14 Carlos Antonio Pinzón , Ehab ElSalamouny , Lucas Massot , Alexis Miller , Héber Hwang Arcolezi , Catuscia Palamidessi

In recent years, neural network-based anomaly detection methods have attracted considerable attention in the hyperspectral remote sensing domain due to the powerful reconstruction ability compared with traditional methods. However, actual…

Computer Vision and Pattern Recognition · Computer Science 2021-05-17 Shaoqi Yu , Xiaorun Li , Shuhan Chen , Liaoying Zhao

Data perturbation-based privacy-preserving methods have been widely adopted in various scenarios due to their efficiency and the elimination of the need for a trusted third party. However, these methods primarily focus on individual…

Cryptography and Security · Computer Science 2025-07-24 Hao Jiang , Quan Zhou , Dongdong Zhao , Shangshang Yang , Wenjian Luo , Xingyi Zhang

Density dependence occurs at the individual level and thus is greatly influenced by spatial local heterogeneity in habitat conditions. However, density dependence is often evaluated at the population level, leading to difficulties or even…

Populations and Evolution · Quantitative Biology 2025-11-20 Qing Zhao , Yunyi Shen

Dataset distillation (DD) has emerged as a widely adopted technique for crafting a synthetic dataset that captures the essential information of a training dataset, facilitating the training of accurate neural models. Its applications span…

Machine Learning · Computer Science 2025-02-04 Saeed Vahidian , Mingyu Wang , Jianyang Gu , Vyacheslav Kungurtsev , Wei Jiang , Yiran Chen

The difficulty of getting medical treatment is one of major livelihood issues in China. Since patients lack prior knowledge about the spatial distribution and the capacity of hospitals, some hospitals have abnormally high or sporadic…

Social and Information Networks · Computer Science 2017-08-03 Hanqing Chao , Yuan Cao , Junping Zhang , Fen Xia , Ye Zhou , Hongming Shan

While non-invasive sampling is more and more commonly used in capture-recapture (CR) experiments, it carries a higher risk of misidentifications than direct observations. As a consequence, one must screen the data to retain only the…

Quantitative Methods · Quantitative Biology 2023-04-04 Rémi Fraysse , Rémi Choquet , Carlo Costantini , Roger Pradel

Subsampling from a large data set is useful in many supervised learning contexts to provide a global view of the data based on only a fraction of the observations. Diverse (or space-filling) subsampling is an appealing subsampling approach…

Methodology · Statistics 2023-11-27 Boyang Shang , Daniel W. Apley , Sanjay Mehrotra

Collecting complete network data is expensive, time-consuming, and often infeasible. Aggregated Relational Data (ARD), which capture information about a social network by asking a respondent questions of the form ``How many people with…

Methodology · Statistics 2022-10-24 Emily Breza , Arun G. Chandrasekhar , Shane Lubold , Tyler H. McCormick , Mengjie Pan