English
Related papers

Related papers: Shaping the Prior: How Synthetic Task Distribution…

200 papers

Recent advances in generative models facilitate the creation of synthetic data to be made available for research in privacy-sensitive contexts. However, the analysis of synthetic data raises a unique set of methodological challenges. In…

Generative modelling has become the standard approach for synthesising tabular data. However, different use cases demand synthetic data to comply with different requirements to be useful in practice. In this survey, we review deep…

Machine Learning · Computer Science 2026-03-17 Mihaela Cătălina Stoian , Eleonora Giunchiglia , Thomas Lukasiewicz

Offline reinforcement learning (RL) is challenged by the distributional shift problem. To address this problem, existing works mainly focus on designing sophisticated policy constraints between the learned policy and the behavior policy.…

Machine Learning · Computer Science 2025-01-09 Yang Yue , Bingyi Kang , Xiao Ma , Qisen Yang , Gao Huang , Shiji Song , Shuicheng Yan

Foundation models (FMs) are increasingly deployed in open-world settings where distribution shift is the rule rather than the exception. The out-of-distribution (OOD) phenomena they face -- knowledge boundaries, capability ceilings,…

Machine Learning · Computer Science 2026-05-08 Xin Wang , Haibo Chen , Wenxuan Liu , Wenwu Zhu

Generating high-fidelity synthetic tabular data under formal differential privacy guarantees remains an open challenge. Methods that provide strong theoretical protection typically sacrifice the modeling of inter-feature dependencies…

Machine Learning · Computer Science 2026-05-27 M. Youssef , M. Woźniak

Diffusion-based image generators are promising priors for ill-posed inverse problems like sparse-view X-ray Computed Tomography (CT). As most studies consider synthetic data, it is not clear whether training data mismatch (``domain shift'')…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Nelas J. Thomsen , Xinyuan Wang , Felix Lucka , Ezgi Demircan-Tureyen

One of the key factors in language productivity and human cognition is the ability of systematic compositionality, which refers to understanding composed unseen examples of seen primitives. However, recent evidence reveals that the…

Computation and Language · Computer Science 2023-12-13 Chen Huang , Peixin Qin , Wenqiang Lei , Jiancheng Lv

Multiscale spatial structure complicates temporal prediction because small-scale spatial fluctuations influence large-scale evolution, yet resolving all scales is often intractable. Standard diffusion models do not address this problem…

Fluid Dynamics · Physics 2026-04-02 Yuki Yasuda , Tobias Bischoff

Accurate prediction of mechanical properties of steel during hot rolling processes, such as Thin Slab Direct Rolling (TSDR), remains challenging due to complex interactions among chemical compositions, processing parameters, and resultant…

Pre-training produces representations that are effective for a wide range of downstream tasks, but it is still unclear what properties of pre-training are necessary for effective gains. Notably, recent work shows that even pre-training on…

Machine Learning · Computer Science 2022-06-22 Yuhuai Wu , Felix Li , Percy Liang

The ability to generate synthetic data has a variety of use cases across different domains. In education research, there is a growing need to have access to synthetic data to test certain concepts and ideas. In recent years, several deep…

Machine Learning · Computer Science 2022-10-18 Herkulaas MvE Combrink , Vukosi Marivate , Benjamin Rosman

Synthetic data can improve generalization when real data is scarce, but excessive reliance may introduce distributional mismatches that degrade performance. In this paper, we present a learning-theoretic framework to quantify the trade-off…

Machine Learning · Statistics 2026-04-02 Amitis Shidani , Tyler Farghly , Yang Sun , Habib Ganjgahi , George Deligiannidis

Synthetic data generation is important to training and evaluating neural models for question answering over knowledge graphs. The quality of the data and the partitioning of the datasets into training, validation and test splits impact the…

Information Retrieval · Computer Science 2020-09-11 Trond Linjordet , Krisztian Balog

Score-based diffusion models generate new samples by learning the score function associated with a diffusion process. While the effectiveness of these models can be theoretically explained using differential equations related to the…

Machine Learning · Computer Science 2026-01-21 Doheon Kim

Tabular synthetic data generators are typically trained to match observational distributions, which can yield high conventional utility (e.g., column correlations, predictive accuracy) yet poor preservation of structural relations relevant…

Machine Learning · Computer Science 2026-03-03 Amir Asiaee , Zhuohui J. Liang , Chao Yan

Injecting rule-based models like Random Forests into differentiable neural network frameworks remains an open challenge in machine learning. Recent advancements have demonstrated that pretrained models can generate efficient molecular…

Machine Learning · Computer Science 2025-05-16 Guillaume Godin

Method validation and study design in causal inference rely on synthetic data with known counterfactuals. Existing simulators trade off distributional realism, the ability to capture mixed-type and multimodal tabular data, against causal…

Methodology · Statistics 2026-03-05 Qi Zhang , Harsh Parikh , Ashley Naimi , Razieh Nabi , Christopher Kim , Timothy Lash

Time series forecasting drives operational decisions in areas like finance, transportation, and energy. While supervised learning approaches achieve strong performance, they require domain-specific training, feature engineering, and ongoing…

Machine Learning · Computer Science 2026-05-26 Kavin Soni , Debanshu Das , Vamshi Guduguntla

Mixture transition distribution time series models build high-order dependence through a weighted combination of first-order transition densities for each one of a specified number of lags. We present a framework to construct stationary…

Methodology · Statistics 2025-02-25 Xiaotian Zheng , Athanasios Kottas , Bruno Sansó

We present Primal, a deterministic feature mapping framework that harnesses the number-theoretic independence of prime square roots to construct robust, tunable vector representations. Diverging from standard stochastic projections (e.g.,…

Machine Learning · Computer Science 2025-11-27 Vladimer Khasia