English

Benchmarking Simulacra AI's Quantum Accurate Synthetic Data Generation for Chemical Sciences

Chemical Physics 2025-11-12 v1 Artificial Intelligence Computational Physics

Abstract

In this work, we benchmark \simulacra's synthetic data generation pipeline against a state-of-the-art Microsoft pipeline on a dataset of small to large systems. By analyzing the energy quality, autocorrelation times, and effective sample size, our findings show that Simulacra's Large Wavefunction Models (LWM) pipeline, paired with state-of-the-art Variational Monte Carlo (VMC) sampling algorithms, reduces data generation costs by 15-50x, while maintaining parity in energy accuracy, and 2-3x compared to traditional CCSD methods on the scale of amino acids. This enables the creation of affordable, large-scale \textit{ab-initio} datasets, accelerating AI-driven optimization and discovery in the pharmaceutical industry and beyond. Our improvements are based on a novel and proprietary sampling scheme called Replica Exchange with Langevin Adaptive eXploration (RELAX).

Keywords

Cite

@article{arxiv.2511.07433,
  title  = {Benchmarking Simulacra AI's Quantum Accurate Synthetic Data Generation for Chemical Sciences},
  author = {Fabio Falcioni and Elena Orlova and Timothy Heightman and Philip Mantrov and Aleksei Ustimenko},
  journal= {arXiv preprint arXiv:2511.07433},
  year   = {2025}
}
R2 v1 2026-07-01T07:30:26.652Z