中文
相关论文

相关论文: Who Are We Missing? A Principled Approach to Chara…

200 篇论文

Machine learning models (e.g., speech recognizers) are usually trained to minimize average loss, which results in representation disparity---minority groups (e.g., non-native speakers) contribute less to the training objective and thus tend…

机器学习 · 统计学 2018-08-01 Tatsunori B. Hashimoto , Megha Srivastava , Hongseok Namkoong , Percy Liang

What proportion of treated units actually benefited from an experimental intervention? What is the median or the largest individual treatment effect? This paper develops methods for answering such questions about the distribution of…

统计方法学 · 统计学 2026-05-11 David Kim , Yongchang Su , Jake Bowers , Xinran Li

Many applications of causal inference require using treatment effects estimated on a study population to make decisions in a separate target population. We consider the challenging setting where there are covariates that are observed in the…

机器学习 · 计算机科学 2024-10-22 Khurram Yamin , Vibhhu Sharma , Ed Kennedy , Bryan Wilder

Given thousands of equally accurate machine learning (ML) models, how can users choose among them? A recent ML technique enables domain experts and data scientists to generate a complete Rashomon set for sparse decision trees--a huge set of…

人机交互 · 计算机科学 2023-03-02 Zijie J. Wang , Chudi Zhong , Rui Xin , Takuya Takagi , Zhi Chen , Duen Horng Chau , Cynthia Rudin , Margo Seltzer

We focus on the problem of generalizing a causal effect estimated on a randomized controlled trial (RCT) to a target population described by a set of covariates from observational data. Available methods such as inverse propensity sampling…

统计方法学 · 统计学 2023-02-27 Imke Mayer , Julie Josse , Traumabase Group

Recent work has shown that standard training via empirical risk minimization (ERM) can produce models that achieve high accuracy on average but low accuracy on underrepresented groups due to the prevalence of spurious features. A…

机器学习 · 计算机科学 2023-05-11 Yachuan Liu , Bohan Zhang , Qiaozhu Mei , Paramveer Dhillon

Automated segmentation of large volumes of medical images is often plagued by the limited availability of fully annotated data and the diversity of organ surface properties resulting from the use of different acquisition protocols for…

计算机视觉与模式识别 · 计算机科学 2023-09-12 Yazhou Zhu , Shidong Wang , Tong Xin , Haofeng Zhang

Individualizing treatment assignment can improve outcomes for diseases with patient-to-patient variability in comparative treatment effects. When a clinical trial demonstrates that some patients improve on treatment while others do not, it…

统计方法学 · 统计学 2022-11-02 Nina Galanter , Marco Carone , Ronald C. Kessler , Alex Luedtke

Training data are usually limited or heterogeneous in many chemical and biological applications. Existing machine learning models for chemistry and materials science fail to consider generalizing beyond training domains. In this article, we…

机器学习 · 计算机科学 2023-10-31 Fang Wu , Nicolas Courty , Shuting Jin , Stan Z. Li

Cluster randomized trials (CRTs) often enroll large numbers of participants, but due to logistical and fiscal challenges, only a subset of participants may be selected for measurement of certain outcomes, and those sampled may, purposely or…

统计方法学 · 统计学 2023-05-16 Joshua R. Nugent , Carina Marquez , Edwin D. Charlebois , Rachel Abbott , Laura B. Balzer

Deep representation learning is a subfield of machine learning that focuses on learning meaningful and useful representations of data through deep neural networks. However, existing methods for semantic classification typically employ…

计算机视觉与模式识别 · 计算机科学 2024-08-30 Kangjun Liu , Ke Chen , Kui Jia , Yaowei Wang

Analyses of cluster randomized trials (CRTs) can be complicated by informative missing outcome data. Methods such as inverse probability weighted generalized estimating equations have been proposed to account for informative missingness by…

统计方法学 · 统计学 2023-04-13 Chia-Rui Chang , Rui Wang

Stepped wedge cluster-randomized trial (CRTs) designs randomize clusters of individuals to intervention sequences, ensuring that every cluster eventually transitions from a control period to receive the intervention under study by the end…

统计方法学 · 统计学 2025-02-19 Alessandro Gasparini , Michael J. Crowther , Emiel O. Hoogendijk , Fan Li , Michael O. Harhay

In our data world, a host of not necessarily trusted controllers gather data on individual subjects. To preserve her privacy and, more generally, her informational self-determination, the individual has to be empowered by giving her agency…

密码学与安全 · 计算机科学 2020-12-22 Josep Domingo-Ferrer , Jordi Soria-Comas

In biomedical research, to obtain more accurate prediction results from a target study, leveraging information from multiple similar source studies is proved to be useful. However, in many biomedical applications based on real-world data,…

统计方法学 · 统计学 2025-12-29 Xiaokang Liu , Jie Hu , Naimin Jing , Yang Ning , Cheng Yong Tang , Runze Li , Yong Chen

We present an optimized rerandomization design procedure for a non-sequential treatment-control experiment. Randomized experiments are the gold standard for finding causal effects in nature. But sometimes random assignments result in…

统计方法学 · 统计学 2021-01-26 Adam Kapelner , Abba M. Krieger , Michael Sklar , David Azriel

Precise estimation of treatment effects is crucial for evaluating intervention effectiveness. While deep learning models have exhibited promising performance in learning counterfactual representations for treatment effect estimation (TEE),…

机器学习 · 计算机科学 2024-01-24 Seungyeon Lee , Ruoqi Liu , Wenyu Song , Lang Li , Ping Zhang

Background: Several studies show that large language models (LLMs) struggle with phenotype-driven gene prioritization for rare diseases. These studies typically use Human Phenotype Ontology (HPO) terms to prompt foundation models like GPT…

计算与语言 · 计算机科学 2026-02-19 Zhanliang Wang , Da Wu , Quan Nguyen , Kai Wang

Subpopulation shift exists widely in many real-world applications, which refers to the training and test distributions that contain the same subpopulation groups but with different subpopulation proportions. Ignoring subpopulation shifts…

In political campaigning substantial resources are spent on voter mobilization, that is, on identifying and influencing as many people as possible to vote. Campaigns use statistical tools for deciding whom to target ("microtargeting"). In…

应用统计 · 统计学 2013-12-02 Thomas Rusch , Ilro Lee , Kurt Hornik , Wolfgang Jank , Achim Zeileis