NASA and IBM release open source Lunar Foundation Model trained on 17 years of orbiter data

NASA and IBM release open source Lunar Foundation Model trained on 17 years of orbiter data

NASA and IBM Research, working with several academic institutions, have released the NASA-IBM Lunar Foundation Model. The two organizations describe it as one of the first open-source foundation models for lunar science. Kevin Murphy, NASA's chief science data officer, says NASA has spent decades building an extraordinary scientific record of the Moon, but collecting data is only part of the job; the data also has to be easier for scientists to use.

A foundation model is pretrained on large volumes of unlabeled data and then adapted to specific tasks with just a few labeled examples. The team sees that as the main benefit for lunar research, where observation data is plentiful but labels are scarce.

The model was trained from scratch on SomBench, which the team says is the largest co-registered multimodal lunar corpus to date. It holds nearly 2 million tile bundles across 11 modalities and two spatial scales. About 1 million high-resolution images come from the Lunar Reconnaissance Orbiter's Narrow Angle Camera at roughly 1 meter per pixel, and just under 964,000 multispectral images come from the Wide Angle Camera at 100 meters per pixel. The bulk of the data is 17 years of LRO observations, which NASA says exceed the data volume of all other NASA planetary missions combined. GRAIL (gravity field), Lunar Prospector (hydrogen) and JAXA's Kaguya/SELENE probe (mineralogy) fill out the collection. In total the dataset has more than 30 spatially aligned data layers from nine instruments and four missions. To prevent leakage between training, validation and test data, the corpus is split geographically by map zones rather than by randomly distributing tiles.

The model is based on TerraMind, a multimodal Earth observation model that IBM developed in 2025 with ESA and Forschungszentrum Jülich, but the researchers trained it from scratch instead of fine-tuning it. For each tile it receives the imaging geometry as explicit context: illumination angles, sun position and tile extent. On the Moon, lighting geometry shapes how the surface looks far more than the surface's actual properties do, so the team feeds the model that information rather than forcing it to correct for lighting from raw pixels. It also learns from high-resolution and coarse imagery together in one training run, and a technique called FlexiViT lets the same trained model handle different image patch sizes without retraining.

The tests covered crater detection at 100-meter and 1-meter scales, prediction of polar ice deposits, and segmentation of Irregular Mare Patches (IMPs), young volcanic features that challenge established timelines of lunar cooling. According to the technical report, the pretrained model matched or beat both common baselines and an architecturally identical control model with random initialization across all tasks.

The biggest gains came in ice deposit prediction. Permanently shadowed polar regions stay cold enough to preserve water ice for billions of years and are considered a potential resource for water, oxygen and rocket fuel. According to IBM, the model cut prediction error by up to 22 percent against the best baseline, SwinV2-B. For coarse-scale crater detection, IBM says the model beat SwinV2-B by nearly 19 percent, a figure that comes from training with only half the data. On meter-scale crater detection and IMP segmentation, the model roughly ties the strongest baselines, with differences inside the variance between training runs. IBM claims a 3 percent lead over SwinV2-B on IMPs, but the article says the results look more comparable than clearly better.

The researchers say part of the edge comes from how the model handles data: it gives each data layer its own processing path, whereas baselines treat all inputs as stacked channels. Even the randomly initialized control model beat five of seven baselines on ice prediction without any lunar pretraining. LoRA, a lighter fine-tuning method that trains only a fraction of the parameters, kept pace with full fine-tuning across the board and did better on crater detection. Full fine-tuning wins only on the two smallest tasks.

Juan Bernabé-Moreno, director of IBM Research Europe, said the model connects observations across instruments and reveals patterns that are difficult to spot in isolation. The limits are explicit. The model isn't suited for absolute geodetic positioning: in generation tests, latitude and longitude were off by dozens of degrees in some cases, and elevation structures could appear with shifted absolute height values even when their shapes were reconstructed correctly. The authors see it as a reusable foundation for downstream tasks, not a replacement for physical measurement instruments. Controlled experiments isolating the contribution of each innovation are still pending, and some test datasets are small.

The model is publicly available on Hugging Face, the code is on GitHub, and it is integrated into the open-source toolkit TerraTorch. The team also released the ML-ready pretraining datasets and benchmark collections. It is part of the NASA-IBM "AI for Science" collaboration; the two organizations have worked on foundation models under a Space Act Agreement since early 2022, and in August 2023 released the first Prithvi model on Hugging Face, trained on Landsat and Sentinel-2 imagery of the contiguous US and adapted for flood and wildfire mapping.

Key facts

  • NASA and IBM Research, with several academic institutions, released the open-source NASA-IBM Lunar Foundation Model, which they describe as one of the first open-source foundation models for lunar science.
  • It was trained from scratch on SomBench: nearly 2 million tile bundles across 11 modalities, more than 30 aligned data layers, nine instruments and four missions, mostly from 17 years of Lunar Reconnaissance Orbiter data.
  • Biggest gain is polar ice deposit prediction, where IBM says error fell by up to 22 percent against the best baseline, SwinV2-B; on meter-scale crater detection and IMP segmentation it roughly ties the strongest baselines.
  • The model is not suited for absolute geodetic positioning: latitude and longitude were off by dozens of degrees in some generation tests.
  • Weights are on Hugging Face, code on GitHub, and the model is integrated into TerraTorch, along with the pretraining datasets and benchmark collections.

Why it matters

Lunar observation data is plentiful but labels are scarce, which is the situation foundation models are built for: pretrain on unlabeled data, then adapt with a few labeled examples. This release applies that approach to the Moon at scale, using SomBench, which the team says is the largest co-registered multimodal lunar corpus to date. The headline result is in polar ice prediction, a topic tied to the Moon's permanently shadowed regions, which are considered a potential resource for water, oxygen and rocket fuel. It also extends a NASA-IBM line of work that began under a Space Act Agreement in early 2022 and produced the Prithvi model in August 2023.

Who it affects

Lunar scientists and machine learning researchers are the direct audience: the model, code, ML-ready pretraining datasets and benchmark collections are all public. Teams working on crater detection, polar ice mapping or volcanic features such as Irregular Mare Patches can use it as a starting point instead of training from zero. Researchers in Earth observation may also watch it, since it builds on TerraMind, which IBM developed in 2025 with ESA and Forschungszentrum Jülich.

How to use it

The model is publicly available on Hugging Face, the code is on GitHub, and it is integrated into the open-source toolkit TerraTorch. The pretraining datasets and benchmark collections are released too. According to the researchers, LoRA, which trains only a fraction of the parameters, kept pace with full fine-tuning across the board and did better on crater detection; full fine-tuning won only on the two smallest tasks. FlexiViT lets the same trained model work with different image patch sizes without retraining.

How solid is it

The performance claims come from the technical report and, for the percentage margins, from IBM. The 22 percent, nearly 19 percent and 3 percent figures are IBM-reported comparisons against SwinV2-B, and the source does not state the metric behind them. The 19 percent figure comes from training with only half the data. The tests span crater detection at two scales, polar ice prediction and IMP segmentation. Across all of them the pretrained model matched or beat the baselines and a randomly initialized control. Even that control beat five of seven baselines on ice prediction, so part of the gain comes from architecture rather than lunar pretraining.

Risks and caveats

The model is not suited for absolute geodetic positioning. In generation tests, latitude and longitude were off by dozens of degrees in some cases, and elevation structures could appear with shifted absolute height values even when their shapes were right. On meter-scale crater detection and IMP segmentation the model only roughly ties the strongest baselines, with differences inside the variance between training runs, so IBM's 3 percent IMP lead looks more comparable than clearly better. Controlled experiments isolating the contribution of each innovation are still pending, and some test datasets are small. The authors frame it as a reusable foundation for downstream tasks, not a replacement for physical measurement instruments.

“NASA has spent decades building an extraordinary scientific record of the Moon, but collecting data is only part of the job”

— Kevin Murphy, NASA's chief science data officer