English

Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding

Artificial Intelligence 2026-03-10 v1

Abstract

We present DeepEarth, a self-supervised multi-modal world model with Earth4D, a novel planetary-scale 4D space-time positional encoder. Earth4D extends 3D multi-resolution hash encoding to include time, efficiently scaling across the planet over centuries with sub-meter, sub-second precision. Multi-modal encoders (e.g. vision-language models) are fused with Earth4D embeddings and trained via masked reconstruction. We demonstrate Earth4D's expressive power by achieving state-of-the-art performance on an ecological forecasting benchmark. Earth4D with learnable hash probing surpasses a multi-modal foundation model pre-trained on substantially more data. Access open source code and download models at: https://github.com/legel/deepearth

Keywords

Cite

@article{arxiv.2603.07039,
  title  = {Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding},
  author = {Lance Legel and Qin Huang and Brandon Voelker and Daniel Neamati and Patrick Alan Johnson and Favyen Bastani and Jeff Rose and James Ryan Hennessy and Robert Guralnick and Douglas Soltis and Pamela Soltis and Shaowen Wang},
  journal= {arXiv preprint arXiv:2603.07039},
  year   = {2026}
}

Comments

8 pages, 5 figures, 1 table. Presented at 2026 World Modeling Workshop, Mila Quebec

R2 v1 2026-07-01T11:08:15.312Z