IBM and NASA are releasing an AI model trained on decades of observations of the Moon, and making it available openly. The NASA-IBM Lunar Foundation Model is arriving on Hugging Face alongside what its creators describe as the first unified, machine-learning-ready dataset of the Moon.
The problem is not a lack of data; it is the opposite. Instruments have been observing the Moon for decades, producing petabytes of information. Scientists, however, still often work through maps and images manually or build separate models for individual tasks.
Neither approach works well at scale. Going through the data by hand takes too long, while models built for a single task are often low resolution and costly to run. Much of the lunar archive therefore remains difficult to use.
The dataset could end up being more important than the model itself. IBM and NASA scientists combined more than 30 spatially aligned layers from nine instruments across four missions, using data from NASA’s Lunar Reconnaissance Orbiter and GRAIL missions, as well as Japan’s SELENE/Kaguya mission.
There was no comparable public dataset before. Lunar observations have traditionally been stored in different formats and at different resolutions, making them difficult to combine. That has been one of the main barriers to applying machine learning to a field that has accumulated far more observations than researchers can realistically process.










