English
Related papers

Related papers: Data integration using covariate summaries from ex…

200 papers

Data from both a randomized trial and an observational study are sometimes simultaneously available for evaluating the effect of an intervention. The randomized data typically allows for reliable estimation of average treatment effects but…

Methodology · Statistics 2021-12-01 David Cheng , Tianxi Cai

In recent years, rather than enclosing data within a single organization, exchanging and combining data from different domains has become an emerging practice. Many studies have discussed the economic and utility value of data and data…

Computers and Society · Computer Science 2020-12-23 Teruaki Hayashi , Hiroki Sakaji , Hiroyasu Matsushima , Yoshiaki Fukami , Takumi Shimizu , Yukio Ohsawa

Obtaining causally interpretable meta-analysis results is challenging when there are differences in the distribution of effect modifiers between eligible trials. To overcome this, recent work on transportability methods has considered…

Methodology · Statistics 2025-03-10 Tat-Thang Vo , Tran Trong Khoi Le , Sivem Afach , Stijn Vansteelandt

The design of data-driven formulations for machine learning and decision-making with good out-of-sample performance is a key challenge. The observation that good in-sample performance does not guarantee good out-of-sample performance is…

Machine Learning · Statistics 2025-02-04 Amine Bennouna , Bart Van Parys , Ryan Lucas

Statistical integration of diverse data sources is an essential step in the building of generalizable prediction tools, especially in precision health. The invariant features model is a new paradigm for multi-source data integration which…

Methodology · Statistics 2025-03-05 Parker Knight , Ndey Isatou Jobe , Rui Duan

The amount of digitally available but heterogeneous information about the world is remarkable, and new technologies such as self-driving cars, smart homes, or the internet of things may further increase it. In this paper we present…

Artificial Intelligence · Computer Science 2018-03-14 Philipp Geiger , Katja Hofmann , Bernhard Schölkopf

Packets originated from an information source in the network can be highly correlated. These packets are often routed through different paths, and compressing them requires to process them individually. Traditional universal compression…

Information Theory · Computer Science 2019-01-14 Ahmad Beirami , Faramarz Fekri

Federated or multi-site studies have distinct advantages over single-site studies, including increased generalizability, the ability to study underrepresented populations, and the opportunity to study rare exposures and outcomes. However,…

Machine Learning · Statistics 2023-09-25 Larry Han , Zhu Shen , Jose Zubizarreta

This paper studies policy evaluation with multiple data sources, especially in scenarios that involve one experimental dataset with two arms, complemented by a historical dataset generated under a single control arm. We propose novel data…

Machine Learning · Statistics 2024-06-04 Ting Li , Chengchun Shi , Qianglin Wen , Yang Sui , Yongli Qin , Chunbo Lai , Hongtu Zhu

Understanding the nature of high-quality summaries is crucial to further improve the performance of multi-document summarization. We propose an approach to characterize human-written summaries using partial information decomposition, which…

Computation and Language · Computer Science 2024-05-24 Laura Mascarell , Yan L'Homme , Majed El Helou

With the increasing computational power of current supercomputers, the size of data produced by scientific simulations is rapidly growing. To reduce the storage footprint and facilitate scalable post-hoc analyses of such scientific data…

Machine Learning · Computer Science 2021-04-14 Subhashis Hazarika , Ayan Biswas , Phillip J. Wolfram , Earl Lawrence , Nathan Urban

When estimating causal effects, it is important to assess external validity, i.e., determine how useful a given study is to inform a practical question for a specific target population. One challenge is that the covariate distribution in…

Methodology · Statistics 2025-01-03 Zhenghao Zeng , Edward H. Kennedy , Lisa M. Bodnar , Ashley I. Naimi

Throughout the different phases of a drug development program, randomized trials are used to establish the tolerability, safety, and efficacy of a candidate drug. At each stage one aims to optimize the design of future studies by…

Applications · Statistics 2021-02-08 Sebastian Weber , Andrew Gelman , Daniel Lee , Michael Betancourt , Aki Vehtari , Amy Racine

We present arguments for the formulation of unified approach to different standard continuous inference methods from partial information. It is claimed that an explicit partition of information into a priori (prior knowledge) and a…

Machine Learning · Statistics 2012-12-07 Mark A. Kon , Leszek Plaskota

In data fusion analysts seek to combine information from two databases comprised of disjoint sets of individuals, in which some variables appear in both databases and other variables appear in only one database. Most data fusion techniques…

Methodology · Statistics 2015-06-22 Bailey K. Fosdick , Maria DeYoreo , Jerome P. Reiter

With the recent developments in digitisation, there are increasing number of documents available online. There are several information extraction tools that are available to extract information from digitised documents. However, identifying…

Information Retrieval · Computer Science 2021-11-08 Richi Nayak , Thirunavukarasu Balasubramaniam , Sangeetha Kutty , Sachindra Banduthilaka , Erin Peterson

We are interested in estimating the effect of a treatment applied to individuals at multiple sites, where data is stored locally for each site. Due to privacy constraints, individual-level data cannot be shared across sites; the sites may…

Machine Learning · Computer Science 2023-04-04 Ruoxuan Xiong , Allison Koenecke , Michael Powell , Zhu Shen , Joshua T. Vogelstein , Susan Athey

A major challenge in nuclear fusion research is the coherent combination of data from heterogeneous diagnostics and modelling codes for machine control and safety as well as physics studies. Measured data from different diagnostics often…

Meta-analysis is commonly used to combine results from multiple clinical trials, but traditional meta-analysis methods do not refer explicitly to a population of individuals to whom the results apply and it is not clear how to use their…

In the analysis of large/big data sets, aggregation (replacing values of a variable over a group by a single value) is a standard way of reducing the size (complexity) of the data. Data analysis programs provide different aggregation…

Machine Learning · Computer Science 2023-03-29 Vladimir Batagelj
‹ Prev 1 4 5 6 7 8 10 Next ›