中文
相关论文

相关论文: Addressing Census data problems in race imputation…

200 篇论文

Multiple measures, such as WEAT or MAC, attempt to quantify the magnitude of bias present in word embeddings in terms of a single-number metric. However, such metrics and the related statistical significance calculations rely on treating…

计算与语言 · 计算机科学 2023-06-16 Alicja Dobrzeniecka , Rafal Urbaniak

While extremely useful (e.g., for COVID-19 forecasting and policy-making, urban mobility analysis and marketing, and obtaining business insights), location data collected from mobile devices often contain data from a biased population…

机器学习 · 计算机科学 2024-02-20 Sepanta Zeighami , Cyrus Shahabi

In almost any geostatistical analysis, one of the underlying, often implicit, modelling assump- tions is that the spatial locations, where measurements are taken, are recorded without error. In this study we develop geostatistical inference…

应用统计 · 统计学 2017-11-02 Claudio Fronterrè , Emanuele Giorgi , Peter J. Diggle

Bi-clustering is a useful approach in analyzing biological data when observations come from heterogeneous groups and have a large number of features. We outline a general Bayesian approach in tackling bi-clustering problems in moderate to…

应用统计 · 统计学 2021-02-11 Han Yan , Jiexing Wu , Yang Li , Jun S. Liu

Missing data is a common issue in various fields such as medicine, social sciences, and natural sciences, and it poses significant challenges for accurate statistical analysis. Although numerous imputation methods have been proposed to…

统计方法学 · 统计学 2025-07-23 Seongmin Kim , Jeunghun Oh , Hungkuk Ko , Jeongmin Park , Jaeyong Lee

Genetic variation in human populations is influenced by geographic ancestry due to spatial locality in historical mating and migration patterns. Spatial population structure in genetic datasets has been traditionally analyzed using either…

种群与进化 · 定量生物学 2016-10-26 Anand Bhaskar , Adel Javanmard , Thomas A. Courtade , David Tse

Current face recognition systems achieve high progress on several benchmark tests. Despite this progress, recent works showed that these systems are strongly biased against demographic sub-groups. Consequently, an easily integrable solution…

计算机视觉与模式识别 · 计算机科学 2020-11-06 Philipp Terhörst , Jan Niklas Kolf , Naser Damer , Florian Kirchbuchner , Arjan Kuijper

Current face recognition systems robustly recognize identities across a wide variety of imaging conditions. In these systems recognition is performed via classification into known identities obtained from supervised identity annotations.…

计算机视觉与模式识别 · 计算机科学 2018-10-16 Daniel C. Castro , Sebastian Nowozin

Biometric recognition is used across a variety of applications from cyber security to border security. Recent research has focused on ensuring biometric performance (false negatives and false positives) is fair across demographic groups.…

统计方法学 · 统计学 2022-08-24 Michael Schuckers , Sandip Purnapatra , Kaniz Fatima , Daqing Hou , Stephanie Schuckers

An important problem in network science is finding an optimal placement of sensors in nodes in order to uniquely detect failures in the network. This problem can be modelled as an identifying code set (ICS) problem, introduced by Karpovsky…

社会与信息网络 · 计算机科学 2023-06-29 Anna L. D. Latour , Arunabha Sen , Kuldeep S. Meel

Approximate string-matching methods to account for complex variation in highly discriminatory text fields, such as personal names, can enhance probabilistic record linkage. However, discriminating between matching and non-matching strings…

We introduce Bayesian spatial change of support methodology for count-valued survey data with known survey variances. Our proposed methodology is motivated by the American Community Survey (ACS), an ongoing survey administered by the U.S.…

应用统计 · 统计学 2014-10-29 Jonathan R. Bradley , Christopher K. Wikle , Scott H. Holan

Network models provide a powerful framework for analysing single-cell count data, facilitating the characterisation of cellular identities, disease mechanisms, and developmental trajectories. However, uncertainty modeling in unsupervised…

基因组学 · 定量生物学 2026-04-27 Shanshan Ren , Thomas E. Bartlett , Lina Gerontogianni , Swati Chandna

Dataset biases are notoriously detrimental to model robustness and generalization. The identify-emphasize paradigm appears to be effective in dealing with unknown biases. However, we discover that it is still plagued by two challenges: A,…

机器学习 · 计算机科学 2023-02-23 Bowen Zhao , Chen Chen , Qian-Wei Wang , Anfeng He , Shu-Tao Xia

Missing values in covariates due to censoring by signal interference or lack of sensitivity in the measuring devices are common in industrial problems. We propose a full Bayesian solution to the prediction problem with an efficient Markov…

统计方法学 · 统计学 2022-01-21 Caroline Svahn , Mattias Villani

Name-based gender prediction has traditionally categorized individuals as either female or male based on their names, using a binary classification system. That binary approach can be problematic in the cases of gender-neutral names that do…

计算与语言 · 计算机科学 2024-07-09 Zhiwen You , HaeJin Lee , Shubhanshu Mishra , Sullam Jeoung , Apratim Mishra , Jinseok Kim , Jana Diesner

We consider applying Bayesian Variable Selection Regression, or BVSR, to genome-wide association studies and similar large-scale regression problems. Currently, typical genome-wide association studies measure hundreds of thousands, or…

应用统计 · 统计学 2011-10-28 Yongtao Guan , Matthew Stephens

The prevailing method of analyzing GWAS data is still to test each marker individually, although from a statistical point of view it is quite obvious that in case of complex traits such single marker tests are not ideal. Recently several…

应用统计 · 统计学 2015-06-19 Erich Dolejsi , Bernhard Bodenstorfer , Florian Frommlet

Despite of the great efforts during the censuses, occurrence of some nonsampling errors such as coverage error is inevitable. Coverage error which can be classified into two types of under-count and overcount occurs when there is no unique…

应用统计 · 统计学 2019-10-15 Sepideh Mosaferi

The statistical challenges in using big data for making valid statistical inference in the finite population have been well documented in literature. These challenges are due primarily to statistical bias arising from under-coverage in the…

统计方法学 · 统计学 2020-06-19 Jae-kwang Kim , Siu-Ming Tam