Related papers: Cycle-StarNet: Bridging the gap between theory and…
Synthetic data is emerging as a promising solution to the scalability issue of supervised deep learning, especially when real data are difficult to acquire or hard to annotate. Synthetic data generation, however, can itself be prohibitively…
Traditional lost-in-space algorithms, such as those implemented in astrometry.net, solve for spacecraft orientation by matching observed star fields to celestial catalogs using geometric asterisms alone. In this work, we propose a novel…
The observable spectrum of an unresolved binary star system is a superposition of two single-star spectra. Even without a detectable velocity offset between the two stellar components, the combined spectrum of a binary system is in general…
The usefulness of deep learning models in robotics is largely dependent on the availability of training data. Manual annotation of training data is often infeasible. Synthetic data is a viable alternative, but suffers from domain gap. We…
Synthetic training data has gained prominence in numerous learning tasks and scenarios, offering advantages such as dataset augmentation, generalization evaluation, and privacy preservation. Despite these benefits, the efficiency of…
The rapid advancement of generative models, such as Stable Diffusion, raises a key question: how can synthetic data from these models enhance predictive modeling? While they can generate vast amounts of datasets, only a subset meaningfully…
Time-series data augmentation mitigates the issue of insufficient training data for deep learning models. Yet, existing augmentation methods are mainly designed for classification, where class labels can be preserved even if augmentation…
We present a technique to synthesise telluric absorption and emission features both for in-situ wavelength calibration and for their removal from astronomical spectra. While the presented technique is applicable for a wide variety of…
Synthetic data offers the promise of cheap and bountiful training data for settings where labeled real-world data is scarce. However, models trained on synthetic data significantly underperform when evaluated on real-world data. In this…
Synthetic samples from diffusion models are promising for leveraging in training discriminative models as replications of real training datasets. However, we found that the synthetic datasets degrade classification performance over real…
In this study, the fundamental stellar atmospheric parameters (Teff, log g, [Fe/H] and [{\alpha}/Fe]) were derived for low-resolution spectroscopy from LAMOST DR5 with Generative Spectrum Networks (GSN). This follows the same scheme as a…
This paper introduces a novel synthetic dataset that captures urban scenes under a variety of weather conditions, providing pixel-perfect, ground-truth-aligned images to facilitate effective feature alignment across domains. Additionally,…
Cloud segmentation plays a crucial role in image analysis for climate modeling. Manually labeling the training data for cloud segmentation is time-consuming and error-prone. We explore to train segmentation networks with synthetic data due…
Planetary studies demand precise and accurate stellar parameters as input to infer the planetary properties. Different methods often provide different results that could lead to biases in the planetary parameters. In this work, we present a…
Accurate model stellar fluxes are key for the analysis of observations of individual stars or stellar populations. Model spectra differ from real stellar spectra due to limitations of the input physical data and adopted simplifications, but…
Modern computer vision systems increasingly encounter performance limitations in data-scarce domains, where collecting large-scale, high-quality labeled data is costly or impractical. While controllable diffusion models enable scalable…
We study, from an empirical standpoint, the efficacy of synthetic data in real-world scenarios. Leveraging synthetic data for training perception models has become a key strategy embraced by the community due to its efficiency, scalability,…
Accurate and comprehensive clinical documentation is crucial for delivering high-quality healthcare, facilitating effective communication among providers, and ensuring compliance with regulatory requirements. However, manual transcription…
Depth completion aims to predict a dense depth map from a sparse depth input. The acquisition of dense ground truth annotations for depth completion settings can be difficult and, at the same time, a significant domain gap between real…
Recently, increasing attention has been drawn to training semantic segmentation models using synthetic data and computer-generated annotation. However, domain gap remains a major barrier and prevents models learned from synthetic data from…