English
Related papers

Related papers: Research Note: Bayesian Record Linkage with Applic…

200 papers

In the computational study of political redistricting, feasibility necessitates the use of a discretization of regions such as states, counties, and towns. In nearly all cases, researchers use a dual graph, whose vertices represent small…

Discrete Mathematics · Computer Science 2026-04-08 Sara Anderson , Sarah Cannon , Brooke Feinberg , Anne Friedman

We linked names and contact information to publicly available profiles in the Personal Genome Project. These profiles contain medical and genomic information, including details about medications, procedures and diseases, and demographic…

Computers and Society · Computer Science 2013-04-30 Latanya Sweeney , Akua Abu , Julia Winn

Sampling geographically dispersed minority populations poses substantial challenges when individual group membership cannot be directly observed. Although stratified sampling can offer efficiency gains, these gains are typically modest…

Applications · Statistics 2026-05-08 Kyla Chasalow , Eitan Hersh , Kosuke Imai , Laura Royden

Unprecedented human mobility has driven the rapid urbanization around the world. In China, the fraction of population dwelling in cities increased from 17.9% to 52.6% between 1978 and 2012. Such large-scale migration poses challenges for…

Computers and Society · Computer Science 2017-11-27 Yang Yang , Chenhao Tan , Zongtao Liu , Fei Wu , Yueting Zhuang

Probabilistic record linkage is often used to match records from two files, in particular when the variables common to both files comprise imperfectly measured identifiers like names and demographic variables. We consider bipartite record…

Methodology · Statistics 2023-12-06 Eric A. Bai , Olivier Binette , Jerome P. Reiter

Despite a large body of literature on trip inference using call detail record (CDR) data, a fundamental understanding of their limitations is lacking. In particular, because of the sparse nature of CDR data, users may travel to a location…

Applications · Statistics 2023-10-09 Zhan Zhao , Haris N. Koutsopoulos , Jinhua Zhao

Anonymized electronic medical records are an increasingly popular source of research data. However, these datasets often lack race and ethnicity information. This creates problems for researchers modeling human disease, as race and…

Quantitative Methods · Quantitative Biology 2018-05-01 Ji-Sung Kim , Xin Gao , Andrey Rzhetsky

Population migration is valuable information which leads to proper decision in urban-planning strategy, massive investment, and many other fields. For instance, inter-city migration is a posterior evidence to see if the government's…

Machine Learning · Statistics 2016-04-08 Renyu Zhao

Biographical databases contain diverse information about individuals. Person names, birth information, career, friends, family and special achievements are some possible items in the record for an individual. The relationships between…

Digital Libraries · Computer Science 2017-09-12 Chao-Lin Liu , Hongsu Wang

Similarity search is a fundamental building block for information retrieval on a variety of datasets. The notion of a neighbor is often based on binary considerations, such as the k nearest neighbors. However, considering that data is often…

Information Retrieval · Computer Science 2022-08-23 Cole Foster , Berk Sevilmis , Benjamin Kimia

Aggregated relative search frequencies offer a unique composite signal reflecting people's habits, concerns, interests, intents, and general information needs, which are not found in other readily available datasets. Temporal search trends…

The difficulty of getting medical treatment is one of major livelihood issues in China. Since patients lack prior knowledge about the spatial distribution and the capacity of hospitals, some hospitals have abnormally high or sporadic…

Social and Information Networks · Computer Science 2017-08-03 Hanqing Chao , Yuan Cao , Junping Zhang , Fen Xia , Ye Zhou , Hongming Shan

Sparse annotation poses persistent challenges to training dense retrieval models; for example, it distorts the training signal when unlabeled relevant documents are used spuriously as negatives in contrastive learning. To alleviate this…

Information Retrieval · Computer Science 2023-10-24 George Zerveas , Navid Rekabsaz , Carsten Eickhoff

Many analyses require linking records from two databases comprising overlapping sets of individuals. In the absence of unique identifiers, the linkage procedure often involves matching on a set of categorical variables, such as…

Applications · Statistics 2017-06-12 Nicole M. Dalzell , Jerome P. Reiter

The identification of urban mobility patterns is very important for predicting and controlling spatial events. In this study, we analyzed millions of geographical check-ins crawled from a leading Chinese location-based social networking…

Physics and Society · Physics 2017-01-03 Zimo Yang , Defu Lian , Nicholas Jing Yuan , Xing Xie , Yong Rui , Tao Zhou

With the need of fast retrieval speed and small memory footprint, document hashing has been playing a crucial role in large-scale information retrieval. To generate high-quality hashing code, both semantics and neighborhood information are…

Information Retrieval · Computer Science 2021-05-28 Zijing Ou , Qinliang Su , Jianxing Yu , Bang Liu , Jingwen Wang , Ruihui Zhao , Changyou Chen , Yefeng Zheng

Many countries today have "country-centric mobile apps" which are mobile apps that are primarily used by residents of a specific country. Many of these country-centric apps also include a location-based service which takes advantage of the…

Social and Information Networks · Computer Science 2019-09-09 Minhui Xue , Xin Yuan , Heather Lee , Keith Ross

Understanding how housing values evolve over time is important to policy makers, consumers and real estate professionals. Existing methods for constructing housing indices are computed at a coarse spatial granularity, such as metropolitan…

Applications · Statistics 2015-05-07 You Ren , Emily B. Fox , Andrew Bruce

In many scenarios, the observational data needed for causal inferences are spread over two data files. In particular, we consider scenarios where one file includes covariates and the treatment measured on one set of individuals, and a…

Methodology · Statistics 2020-09-22 Sharmistha Guha , Jerome P. Reiter , Andrea Mercatanti

In many healthcare and social science applications, information about units is dispersed across multiple data files. Linking records across files is necessary to estimate the associations of interest. Common record linkage algorithms only…

Methodology · Statistics 2024-06-25 Gauri Kamat , Mingyang Shan , Roee Gutman