English
Related papers

Related papers: Data integration using covariate summaries from ex…

200 papers

Data integration methods aim to extract low-dimensional embeddings from high-dimensional outcomes to remove unwanted variations, such as batch effects and unmeasured covariates, across heterogeneous datasets. However, multiple hypothesis…

Methodology · Statistics 2025-12-15 Jin-Hong Du , Kathryn Roeder , Larry Wasserman

One approach for increasing the efficiency of randomized trials is the use of "external controls" -- individuals who received the control treatment studied in the trial during routine practice or in prior experimental studies. Existing…

In studies that rely on data from electronic health records (EHRs), unstructured text data such as clinical progress notes offer a rich source of information about patient characteristics and care that may be missing from structured data.…

Computation and Language · Computer Science 2024-05-22 Reagan Mozer , Aaron R. Kaufman , Leo A. Celi , Luke Miratrix

Many modern applications collect data that comes in federated spirit, with data kept locally and undisclosed. Till date, most insight into the causal inference requires data to be stored in a central repository. We present a novel framework…

Methodology · Statistics 2021-06-02 Thanh Vinh Vo , Trong Nghia Hoang , Young Lee , Tze-Yun Leong

Automatically evaluating the coherence of summaries is of great significance both to enable cost-efficient summarizer evaluation and as a tool for improving coherence by selecting high-scoring candidate summaries. While many different…

Computation and Language · Computer Science 2022-09-16 Julius Steen , Katja Markert

Estimating treatment effects conditional on observed covariates can improve the ability to tailor treatments to particular individuals. Doing so effectively requires dealing with potential confounding, and also enough data to adequately…

Meta-analysis, because of both logistical convenience and statistical efficiency, is widely popular for synthesizing information on common parameters of interest across multiple studies. We propose developing a generalized meta-analysis…

Methodology · Statistics 2018-11-27 Prosenjit Kundu , Runlong Tang , Nilanjan Chatterjee

In many contexts, we have access to aggregate data, but individual level data is unavailable. For example, medical studies sometimes report only aggregate statistics about disease prevalence because of privacy concerns. Even so, many a time…

Machine Learning · Computer Science 2018-09-18 Sanket Tavarageri , Nag Mani , Anand Ramasubramanian , Jaskiran Kalsi

The development and approval of new treatments generates large volumes of results, such as summaries of efficacy and safety. However, it is commonly overlooked that analyzing clinical study data also produces data in the form of results.…

Computers and Society · Computer Science 2022-04-22 Joana M Barros , Lukas A Widmer , Mark Baillie , Simon Wandel

In this paper, we propose a data collaboration analysis method for distributed datasets. The proposed method is a centralized machine learning while training datasets and models remain distributed over some institutions. Recently, data…

Machine Learning · Computer Science 2019-02-21 Akira Imakura , Tetsuya Sakurai

Compositional data represent a specific family of multivariate data, where the information of interest is contained in the ratios between parts rather than in absolute values of single parts. The analysis of such specific data is…

The need for multimodal data integration arises naturally when multiple complementary sets of features are measured on the same sample. Under a dependent multifactor model, we develop a fully data-driven orchestrated approximate message…

Methodology · Statistics 2026-01-22 Sagnik Nandy , Zongming Ma

Data fusion and transfer learning are rapidly growing fields that enhance model performance for a target population by leveraging other related data sources or tasks. The challenges lie in the various potential heterogeneities between the…

Machine Learning · Statistics 2025-08-19 Jing Wang , HaiYing Wang , Kun Chen

In evidence synthesis, multilevel modeling approaches (MMAs) are commonly employed to combine aggregate data (AD) and individual participant data (IPD). These approaches rely on an aggregate outcome model that is ideally obtained by…

Methodology · Statistics 2026-02-18 Tran Trong Khoi Le , Marie-Felicia Béclin , Sivem Afach , Tat-Thang Vo

Open data platforms such as data.gov or opendata.socrata. com provide a huge amount of valuable information. Their free-for-all nature, the lack of publishing standards and the multitude of domains and authors represented on these platforms…

Databases · Computer Science 2012-05-14 Julian Eberius , Katrin Braunschweig , Maik Thiele , Wolfgang Lehner

Many stochastic optimization problems include chance constraints that enforce constraint satisfaction with a specific probability; however, solving an optimization problem with chance constraints assumes that the solver has access to the…

Optimization and Control · Mathematics 2021-09-21 Joshua Comden , Ahmed S. Zamzam , Andrey Bernstein

In modern biomedical research, it is ubiquitous to have multiple data sets measured on the same set of samples from different views (i.e., multi-view data). For example, in genetic studies, multiple genomic data sets at different molecular…

Methodology · Statistics 2017-03-20 Gen Li , Sungkyu Jung

Heterogeneous data from multiple populations, sub-groups, or sources is often represented as a ``mixture model'' with a single latent class influencing all of the observed covariates. Heterogeneity can be resolved at multiple levels by…

Machine Learning · Computer Science 2024-12-16 Bijan Mazaheri , Chandler Squires , Caroline Uhler

An aggregate data meta-analysis is a statistical method that pools the summary statistics of several selected studies to estimate the outcome of interest. When considering a continuous outcome, typically each study must report the same…

Methodology · Statistics 2022-06-22 Sean McGrath , XiaoFei Zhao , Zhi Zhen Qin , Russell Steele , Andrea Benedetti

Statistical estimation in many contemporary settings involves the acquisition, analysis, and aggregation of datasets from multiple sources, which can have significant differences in character and in value. Due to these variations, the…

Applications · Statistics 2014-12-23 Quentin Berthet , Venkat Chandrasekaran
‹ Prev 1 3 4 5 6 7 10 Next ›