Models

NASA and IBM release lunar AI foundation model

NASA and IBM have released an open-source lunar foundation model trained on 17 years of satellite data, offering scientists a powerful new tool to predict ice deposits and detect craters.

The Decoder1 day agoModels
Image: The Decoder

NASA and IBM Research, alongside academic partners, have launched the NASA-IBM Lunar Foundation Model. Built to make decades of lunar observations accessible for machine learning, the open-source model was trained from scratch using SomBench, a multimodal corpus containing nearly 2 million tile bundles across 11 modalities and two spatial scales. This dataset includes roughly 1 million high-resolution images from the Lunar Reconnaissance Orbiter's Narrow Angle Camera at 1 meter per pixel, and nearly 964,000 multispectral images from its Wide Angle Camera at 100 meters per pixel. The training corpus integrates over 30 spatially aligned data layers from nine instruments and four missions, including GRAIL, Lunar Prospector, and JAXA's Kaguya/SELENE probe.

The model adapts the architecture of TerraMind, an Earth observation model developed with ESA and Forschungszentrum Jülich, but was trained from scratch. To handle the Moon's extreme lighting conditions, researchers feed explicit imaging geometry—such as sun position and illumination angles—directly into the model. It processes different data types through individual paths rather than stacking them as channels, and utilizes FlexiViT to accommodate varying image patch sizes without retraining.

In evaluations, the model excelled at predicting polar ice deposits, reducing prediction error by up to 22 percent compared to the SwinV2-B baseline. For coarse-scale crater detection, it outperformed SwinV2-B by nearly 19 percent while using only half the training data. It also achieved a 3 percent lead over the baseline when segmenting Irregular Mare Patches. Practitioners can utilize LoRA for lightweight fine-tuning, which matched or exceeded full fine-tuning on most tasks. However, the model is not designed for absolute geodetic positioning, as generation tests showed latitude and longitude errors of dozens of degrees.

For planetary scientists, this release simplifies the transition from raw observation to active research. Instead of labeling massive datasets from scratch, researchers can adapt this pre-trained model to specific tasks with minimal labeled examples. The model, along with its pretraining datasets, is publicly hosted on Hugging Face and GitHub, and has been integrated into the TerraTorch toolkit.

This is our own summary of reporting by The Decoder

More in Models