English
Related papers

Related papers: Generating Synthetic Text Data to Evaluate Causal …

200 papers

Estimating long-term causal effects by combining long-term observational and short-term experimental data is a crucial but challenging problem in many real-world scenarios. In existing methods, several ideal assumptions, e.g. latent…

Machine Learning · Computer Science 2025-05-12 Ruichu Cai , Junjie Wan , Weilin Chen , Zeqin Yang , Zijian Li , Peng Zhen , Jiecheng Guo

Causal inference is a central goal across many scientific disciplines. Over the past several decades, three major frameworks have emerged to formalize causal questions and guide their analysis: the potential outcomes framework, structural…

Statistics Theory · Mathematics 2026-02-12 Linbo Wang , Thomas Richardson , James Robins

Recent developments in causal machine learning methods have made it easier to estimate flexible relationships between confounders, treatments and outcomes, making unconfoundedness assumptions in causal analysis more palatable. How…

Econometrics · Economics 2026-05-22 Justin Young , Eleanor Wiske Dillon

Recent neural approaches to data-to-text generation have mostly focused on improving content fidelity while lacking explicit control over writing styles (e.g., word choices, sentence structures). More traditional systems use templates to…

Computation and Language · Computer Science 2020-10-12 Shuai Lin , Wentao Wang , Zichao Yang , Xiaodan Liang , Frank F. Xu , Eric Xing , Zhiting Hu

It is generally difficult to make any statements about the expected prediction error in an univariate setting without further knowledge about how the data were generated. Recent work showed that knowledge about the real underlying causal…

Artificial Intelligence · Computer Science 2017-04-18 Patrick Blöbaum , Takashi Washio , Shohei Shimizu

Using observed language to understand interpersonal interactions is important in high-stakes decision making. We propose a causal research design for observational (non-experimental) data to estimate the natural direct and indirect effects…

Computation and Language · Computer Science 2021-09-17 Katherine A. Keith , Douglas Rice , Brendan O'Connor

Making evidence based decisions requires data. However for real-world applications, the privacy of data is critical. Using synthetic data which reflects certain statistical properties of the original data preserves the privacy of the…

Machine Learning · Computer Science 2021-05-28 Varun Chandrasekaran , Darren Edge , Somesh Jha , Amit Sharma , Cheng Zhang , Shruti Tople

Large Language Models (LLMs) generate realistic synthetic data but offer no guarantee that their outputs respect the causal mechanisms governing the target domain. We introduce CausalSynth, a framework that decouples causal structure…

Machine Learning · Computer Science 2026-05-19 Zehua Cheng , Wei Dai , Jiahao Sun , Thomas Lukasiewicz

At the heart of causal structure learning from observational data lies a deceivingly simple question: given two statistically dependent random variables, which one has a causal effect on the other? This is impossible to answer using…

Machine Learning · Computer Science 2020-10-13 Nikolaos Nikolaou , Konstantinos Sechidis

High-quality data is essential for conversational recommendation systems and serves as the cornerstone of the network architecture development and training strategy design. Existing works contribute heavy human efforts to manually labeling…

Computation and Language · Computer Science 2023-06-19 Yu Lu , Junwei Bao , Zichen Ma , Xiaoguang Han , Youzheng Wu , Shuguang Cui , Xiaodong He

This study investigates the consequences of training language models on synthetic data generated by their predecessors, an increasingly prevalent practice given the prominence of powerful generative models. Diverging from the usual emphasis…

Computation and Language · Computer Science 2024-04-17 Yanzhu Guo , Guokan Shang , Michalis Vazirgiannis , Chloé Clavel

In causal inference, it is common to estimate the causal effect of a single treatment variable on an outcome. However, practitioners may also be interested in the effect of simultaneous interventions on multiple covariates of a fixed target…

Methodology · Statistics 2022-11-24 Jaime Roquero Gimenez , Dominik Rothenhäusler

The abundance of data produced daily from large variety of sources has boosted the need of novel approaches on causal inference analysis from observational data. Observational data often contain noisy or missing entries. Moreover, causal…

Methodology · Statistics 2017-03-14 Fani Tsapeli , Peter Tino , Mirco Musolesi

Synthetic data is increasingly critical for contact centers, where privacy constraints and data scarcity limit the availability of real conversations. However, generating synthetic dialogues that are realistic and useful for downstream…

Computation and Language · Computer Science 2026-02-17 Rishikesh Devanathan , Varun Nathan , Ayush Kumar

This paper examines methods of causal inference based on groupwise matching when we observe multiple large groups of individuals over several periods. We formulate causal inference validity through a generalized matching condition,…

Econometrics · Economics 2026-03-24 Ratzanyel Rincón , Kyungchul Song

Large language model (LLM) development is currently driven by large-scale empirical iteration over data mixtures, reward models, routing strategies, and evaluation pipelines. Here, we argue that many central questions in LLM development and…

Curating a large scale medical imaging dataset for machine learning applications is both time consuming and expensive. Balancing the workload between model development, data collection and annotations is difficult for machine learning…

Artificial Intelligence · Computer Science 2022-06-07 Athanasios Vlontzos , Hadrien Reynaud , Bernhard Kainz

From simulating galaxy formation to viral transmission in a pandemic, scientific models play a pivotal role in developing scientific theories and supporting government policy decisions that affect us all. Given these critical applications,…

Software Engineering · Computer Science 2023-07-03 Andrew G. Clark , Michael Foster , Benedikt Prifling , Neil Walkinshaw , Robert M. Hierons , Volker Schmidt , Robert D. Turner

Synthetic tabular data are often evaluated by distributional similarity, privacy distance, or train-on-synthetic-test-on-real predictive performance, but these criteria do not ensure validity for causal inference. We show that fully…

Methodology · Statistics 2026-05-12 Yichen Xu

State-of-the-art models for keyphrase generation require large amounts of training data to achieve good performance. However, obtaining keyphrase-labeled documents can be challenging and costly. To address this issue, we present a…

Computation and Language · Computer Science 2024-11-07 Mael Houbre , Florian Boudin , Beatrice Daille , Akiko Aizawa