中文
相关论文

相关论文: Mechanisms for Data Sharing in Collaborative Causa…

200 篇论文

Traditional causal inference approaches leverage observational study data to estimate the difference in observed and unobserved outcomes for a potential treatment, known as the Conditional Average Treatment Effect (CATE). However, CATE…

机器学习 · 统计学 2022-09-22 Tianhui Zhou , William E Carson , David Carlson

Data scarcity is a tremendous challenge in causal effect estimation. In this paper, we propose to exploit additional data sources to facilitate estimating causal effects in the target population. Specifically, we leverage additional source…

机器学习 · 计算机科学 2021-06-01 Thanh Vinh Vo , Pengfei Wei , Trong Nghia Hoang , Tze-Yun Leong

For a federated learning model to perform well, it is crucial to have a diverse and representative dataset. However, the data contributors may only be concerned with the performance on a specific subset of the population, which may not…

计算机科学与博弈论 · 计算机科学 2023-06-12 Baihe Huang , Sai Praneeth Karimireddy , Michael I. Jordan

Classification is a well-studied machine learning task which concerns the assignment of instances to a set of outcomes. Classification models support the optimization of managerial decision-making across a variety of operational business…

机器学习 · 计算机科学 2025-05-19 Wouter Verbeke , Diego Olaya , Jeroen Berrevoets , Sam Verboven , Sebastián Maldonado

Causal discovery and causal reasoning are classically treated as separate and consecutive tasks: one first infers the causal graph, and then uses it to estimate causal effects of interventions. However, such a two-stage approach is…

Attribution methods are primarily designed to study input component contributions to individual model predictions. However, some research applications require a summary of attribution patterns across the entire dataset to facilitate the…

机器学习 · 计算机科学 2025-07-15 Pierre Lelièvre , Chien-Chung Chen

Causal inference is capable of estimating the treatment effect (i.e., the causal effect of treatment on the outcome) to benefit the decision making in various domains. One fundamental challenge in this research is that the treatment…

机器学习 · 计算机科学 2021-12-28 Qian Li , Zhichao Wang , Shaowu Liu , Gang Li , Guandong Xu

Multi-domain recommendation leverages domain-general knowledge to improve recommendations across several domains. However, as platforms expand to dozens or hundreds of scenarios, training all domains in a unified model leads to performance…

信息检索 · 计算机科学 2025-07-10 Huishi Luo , Yiqing Wu , Yiwen Chen , Fuzhen Zhuang , Deqing Wang

Many of the traditional recommendation algorithms are designed based on the fundamental idea of mining or learning correlative patterns from data to estimate the user-item correlative preference. However, pure correlative learning may lead…

信息检索 · 计算机科学 2023-08-15 Shuyuan Xu , Yingqiang Ge , Yunqi Li , Zuohui Fu , Xu Chen , Yongfeng Zhang

In this paper, we consider a setting where heterogeneous agents with connectivity are performing inference using unlabeled streaming data. Observed data are only partially informative about the target variable of interest. In order to…

机器学习 · 计算机科学 2025-01-28 Mert Kayaalp , Yunus Inan , Visa Koivunen , Ali H. Sayed

Complex systems are ubiquitous in the real world and tend to have complicated and poorly understood dynamics. For their control issues, the challenge is to guarantee accuracy, robustness, and generalization in such bloated and troubled…

人工智能 · 计算机科学 2022-09-16 Xuehui Yu , Jingchi Jiang , Xinmiao Yu , Yi Guan , Xue Li

We propose a framework for adaptive data-centric collaborative machine learning among self-interested agents, coordinated by an arbiter. Designed to handle the incremental nature of real-world data, the framework operates in an online…

机器学习 · 计算机科学 2025-02-07 Nithia Vijayan , Bryan Kian Hsiang Low

Causal representation learning seeks to extract high-level latent factors from low-level sensory data. Most existing methods rely on observational data and structural assumptions (e.g., conditional independence) to identify the latent…

机器学习 · 统计学 2024-02-26 Kartik Ahuja , Divyat Mahajan , Yixin Wang , Yoshua Bengio

Causal inference across multiple data sources offers a promising avenue to enhance the generalizability and replicability of scientific findings. However, data integration methods for time-to-event outcomes, common in biomedical research,…

统计方法学 · 统计学 2025-05-16 Yi Liu , Alexander W. Levis , Ke Zhu , Shu Yang , Peter B. Gilbert , Larry Han

Federated learning promises significant sample-efficiency gains by pooling data across multiple agents, yet incentive misalignment is an obstacle: each update is costly to the contributor but boosts every participant. We introduce a…

计算机科学与博弈论 · 计算机科学 2026-02-02 Ariel D. Procaccia , Han Shao , Itai Shapira

The era of big data has witnessed an increasing availability of observational data from mobile and social networking, online advertising, web mining, healthcare, education, public policy, marketing campaigns, and so on, which facilitates…

机器学习 · 计算机科学 2023-03-06 Zhixuan Chu , Ruopeng Li , Stephen Rathbun , Sheng Li

Classical causal and statistical inference methods typically assume the observed data consists of independent realizations. However, in many applications this assumption is inappropriate due to a network of dependences between units in the…

机器学习 · 计算机科学 2019-07-02 Rohit Bhattacharya , Daniel Malinsky , Ilya Shpitser

Collaboration between different data centers is often challenged by heterogeneity across sites. To account for the heterogeneity, the state-of-the-art method is to re-weight the covariate distributions in each site to match the distribution…

机器学习 · 统计学 2024-04-25 Tianyu Guo , Sai Praneeth Karimireddy , Michael I. Jordan

Cooperative inference across independently deployed machine learning models is increasingly desirable in distributed environments, as there is a growing need to leverage multiple models while keeping their data and model parameters private.…

机器学习 · 计算机科学 2026-05-08 Yui Hashimoto , Takayuki Nishio , Yuichi Kitagawa , Takahito Tanimura

Merging datasets across institutions is a lengthy and costly procedure, especially when it involves private information. Data hosts may therefore want to prospectively gauge which datasets are most beneficial to merge with, without…

机器学习 · 统计学 2025-07-03 Jake Fawkes , Lucile Ter-Minassian , Desi Ivanova , Uri Shalit , Chris Holmes