English
Related papers

Related papers: Bregman Divergence-Based Data Integration with App…

200 papers

Many Mendelian randomization (MR) papers have been conducted only in people of European ancestry, limiting transportability of results to the global population. Expanding MR to diverse ancestry groups is essential to ensure equitable…

Genome-wide association studies (GWAS) have been widely used to examine the association between single nucleotide polymorphisms (SNPs) and complex traits, where both the sample size n and the number of SNPs p can be very large. Recently,…

Methodology · Statistics 2019-03-05 Bingxin Zhao , Hongtu Zhu

The analysis of mixed data has been raising challenges in statistics and machine learning. One of two most prominent challenges is to develop new statistical techniques and methodologies to effectively handle mixed data by making the data…

Machine Learning · Computer Science 2017-08-21 Tu Dinh Nguyen , Truyen Tran , Dinh Phung , Svetha Venkatesh

Estimating racial disparity requires individual-level race data, which are often unavailable due to the sensitivity of collecting such information. To address this problem, many researchers utilize Bayesian Improved Surname Geocoding…

Computation and Language · Computer Science 2026-04-29 Noah Dasanaike , Kosuke Imai

Genome-wide association studies (GWAS) provide a means of examining the common genetic variation underlying a range of traits and disorders. In addition, it is hoped that GWAS may provide a means of differentiating affected from unaffected…

With rapid technological growth, automatic pronunciation assessment has transitioned toward systems that evaluate pronunciation in various aspects, such as fluency and stress. However, despite the highly imbalanced score labels within each…

Computation and Language · Computer Science 2023-08-30 Heejin Do , Yunsu Kim , Gary Geunbae Lee

A key challenge in building effective regression models for large and diverse populations is accounting for patient heterogeneity. An example of such heterogeneity is in health system risk modeling efforts where different combinations of…

Methodology · Statistics 2022-12-26 Jared D. Huling , Menggang Yu

The family of f-divergences is ubiquitously applied to generative modeling in order to adapt the distribution of the model to that of the data. Well-definedness of f-divergences, however, requires the distributions of the data and model to…

Machine Learning · Statistics 2019-06-04 Akash Srivastava , Kristjan Greenewald , Farzaneh Mirzazadeh

Unsupervised person re-identification has achieved great success through the self-improvement of individual neural networks. However, limited by the lack of diversity of discriminant information, a single network has difficulty learning…

Computer Vision and Pattern Recognition · Computer Science 2023-06-09 Yunpeng Zhai , Peixi Peng , Mengxi Jia , Shiyong Li , Weiqiang Chen , Xuesong Gao , Yonghong Tian

Selection bias is a serious potential problem for inference about relationships of scientific interest based on samples without well-defined probability sampling mechanisms. Motivated by the potential for selection bias in (a) estimated…

We propose a simple yet powerful test statistic to quantify the discrepancy between two conditional distributions. The new statistic avoids the explicit estimation of the underlying distributions in highdimensional space and it operates on…

Machine Learning · Computer Science 2021-01-01 Shujian Yu , Ammar Shaker , Francesco Alesiani , Jose C. Principe

Motivated by various applications, we consider the problem of homogeneous human population size (N) estimation from Dual-record system (DRS) (equivalently, two-sample capture-recapture experiment). The likelihood estimate from the…

Methodology · Statistics 2015-04-07 Kiranmoy Chatterjee , Diganta Mukherjee

In the field of healthcare, electronic health records (EHR) serve as crucial training data for developing machine learning models for diagnosis, treatment, and the management of healthcare resources. However, medical datasets are often…

Machine Learning · Statistics 2024-03-21 Keira Behal , Jiayi Chen , Caleb Fikes , Sophia Xiao

In the last decades, it has been discussed the use of epidemiological prevalence ratio (PR) rather than odds ratio as a measure of association to be estimated in cross-sectional studies. The main difficulties in use of statistical models…

The aim of reduced rank regression is to connect multiple response variables to multiple predictors. This model is very popular, especially in biostatistics where multiple measurements on individuals can be re-used to predict multiple…

Methodology · Statistics 2022-06-20 The Tien Mai , Pierre Alquier

Causal inference analyses often use existing observational data, which in many cases has some clustering of individuals. In this paper we discuss propensity score weighting methods in a multilevel setting where within clusters individuals…

Applications · Statistics 2020-12-24 Youjin Lee , Trang Q. Nguyen , Elizabeth A. Stuart

In some socio-economic surveys, data are collected on sensitive or stigmatizing issues such as tax evasion, criminal conviction, drug use, etc. In such surveys, direct questioning of respondents is not of much use and the randomized…

Statistics Theory · Mathematics 2013-03-22 Mausumi Bose

Unbiased estimators are introduced for averaged Bregman divergences which generalize Stein's Unbiased (Predictive) Risk Estimator, and the minimization of these estimators is proposed as a regularization parameter selection method for…

Numerical Analysis · Mathematics 2021-11-22 Elias S. Helou , Sandra A. Santos , Lucas E. A. Simões

This paper develops a sparsity-inducing version of Bayesian Causal Forests, a recently proposed nonparametric causal regression model that employs Bayesian Additive Regression Trees and is specifically designed to estimate heterogeneous…

Methodology · Statistics 2021-11-17 Alberto Caron , Gianluca Baio , Ioanna Manolopoulou

Statistical heterogeneity is a measure of how skewed the samples of a dataset are. It is a common problem in the study of differential privacy that the usage of a statistically heterogeneous dataset results in a significant loss of…

Machine Learning · Computer Science 2024-12-02 Mary Scott , Graham Cormode , Carsten Maple
‹ Prev 1 3 4 5 6 7 10 Next ›