English
Related papers

Related papers: Bregman Divergence-Based Data Integration with App…

200 papers

U.S. state education agencies mark schools displaying achievement gaps between demographic subgroups as needing improvement. Some schools may have few students in these subgroups, such that average end-of-year test scores only noisily…

Methodology · Statistics 2025-12-10 Joshua Wasserman , Michael R. Elliott , Ben B. Hansen

We address the problem of performing regression while ensuring demographic parity, even without access to sensitive attributes during inference. We present a general-purpose post-processing algorithm that, using accurate estimates of the…

Machine Learning · Statistics 2024-07-23 Evgenii Chzhen , Mohamed Hebiri , Gayane Taturyan

Sharing diverse genomic and other biomedical datasets is critical to advance scientific discoveries and their equitable translation to improve human health. However, data sharing remains challenging in the context of legacy datasets,…

Polygenic risk scores (PRSs) can significantly enhance breast cancer risk prediction when combined with clinical risk factor data. While many studies have explored the value-add of PRSs, little is known about the potential impact of…

Genomics · Quantitative Biology 2024-07-31 Monica Isgut , Andrew Hornback , Yunan Luo , Asma Khimani , Neha Jain , May D. Wang

Environmental hazards place certain individuals at disproportionately higher risks. As these hazards increasingly endanger human health, precise identification of the most vulnerable population subgroups is critical for public health.…

Machine Learning · Computer Science 2024-09-23 Jong Woo Nam , Eun Young Choi , Jennifer A. Ailshire , Yao-Yi Chiang

With the development of Big data technology, data analysis has become increasingly important. Traditional clustering algorithms such as K-means are highly sensitive to the initial centroid selection and perform poorly on non-convex…

Machine Learning · Computer Science 2023-07-28 Ying Xiao , Hou-biao Li , Yu-pu Zhang

Estimating the causal treatment effects by subgroups is important in observational studies when the treatment effect heterogeneity may be present. Existing propensity score methods rely on a correctly specified propensity score model. Model…

Methodology · Statistics 2024-04-19 Yan Li , Yong-Fang Kuo , Liang Li

This paper presents an algorithm tailored for the efficient recovery of sparse probability measures incorporating $\ell_0$-sparse regularization within the probability simplex constraint. Employing the Bregman proximal gradient method, our…

Optimization and Control · Mathematics 2024-11-26 Jianting Pan , Ming Yan

This paper introduces an innovative Electronic Health Record (EHR) foundation model that integrates Polygenic Risk Scores (PRS) as a foundational data modality, moving beyond traditional EHR-only approaches to build more holistic health…

The crowdsourcing scenarios are a good example of having a probability distribution over some categories showing what the people in a global perspective thinks. Learn a predictive model of this probability distribution can be of much more…

Machine Learning · Computer Science 2019-01-31 F. A. Mena , R. Ñanculef

The ratio of two probability density functions is a fundamental quantity that appears in many areas of statistics and machine learning, including causal inference, reinforcement learning, covariate shift, outlier detection, independence…

Machine Learning · Statistics 2026-05-13 Oliver J. Hines , Caleb H. Miles

Although Automatic Speech Recognition (ASR) systems have become an integral part of modern technology, their evaluation remains challenging, particularly for low-resource languages such as Persian. This paper introduces Persian Speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-28 Nima Sedghiyeh , Sara Sadeghi , Reza Khodadadi , Farzin Kashani , Omid Aghdaei , Somayeh Rahimi , Mohammad Sadegh Safari

Logarithmic score and information divergence appear in both information theory, statistics, statistical mechanics, and portfolio theory. We demonstrate that all these topics involve some kind of optimization that leads directly to the use…

Statistics Theory · Mathematics 2015-07-28 Peter Harremoës

Differential privacy is a cryptographically-motivated definition of privacy which has gained significant attention over the past few years. Differentially private solutions enforce privacy by adding random noise to a function computed over…

Machine Learning · Computer Science 2012-07-03 Kamalika Chaudhuri , Daniel Hsu

In heterogeneous cohorts and those where censoring by non-primary risks is informative many conventional survival analysis methods are not applicable; the proportional hazards assumption is usually violated at population level and the…

Bipolar Disorder (BD) is a complex disease. It is heterogeneous, both at the phenotypic and genetic level, although the extent and impact of this heterogeneity is not fully understood. In this paper, we leverage recent advances in…

Preserving privacy and reducing communication costs for edge users pose significant challenges in recommendation systems. Although federated learning has proven effective in protecting privacy by avoiding data exchange between clients and…

Machine Learning · Computer Science 2023-11-01 Lin Wang , Zhichao Wang , Xi Leng , Xiaoying Tang

Risk for autism can be influenced by genetic mutations in hundreds of genes. Based on findings showing that genes with highly correlated gene expressions are functionally interrelated, "guilt by association" methods such as DAWN have been…

Methodology · Statistics 2020-03-10 Kevin Z. Lin , Han Liu , Kathryn Roeder

Propensity score weighting approaches have been widely implemented in clinical research to estimate the effects of a treatment or exposure while mitigating the risk of confounding in the absence of random assignment. In practice, when…

Methodology · Statistics 2026-04-17 Emma K. Mackay , Amol A. Verma , Fahad Razak , Surain B. Roberts

Labeled speech data from patients with Parkinsons disease (PD) are scarce, and the statistical distributions of training and test data differ significantly in the existing datasets. To solve these problems, dimensional reduction and sample…

Machine Learning · Computer Science 2020-02-11 Xiaoheng Zhang , Yongming Li , Pin Wang , Xiaoheng Tan , Yuchuan Liu