IBM and NASA Release Open-Source AI Model for Lunar Exploration

Patterns that are difficult to see in isolation
IBM's research director on what the model reveals by connecting observations across multiple lunar instruments.
Mark

Why does it matter that this is open-source? Couldn't NASA just keep this tool for itself?

Mimi

Open-source means any researcher anywhere can use it, modify it, and build on it. If NASA kept it proprietary, only their teams could benefit. By releasing it openly, they're saying: here's a foundation—take it and solve problems we haven't even thought of yet.

Luke

But I want to be clear on what "open-source" actually means here. Is the model itself freely available for download? Is the training code available? The article says the model and dataset are released, but it doesn't specify the license or whether there are computational barriers to actually using it.

Mark

The article mentions this reduces error by 22 percent on ice detection. How significant is that improvement?

Mimi

In lunar science, a 22 percent reduction in error when you're looking for ice deposits is substantial. Ice indicates water and oxygen—resources you need to sustain human presence on the Moon. A more accurate model means fewer false leads, less wasted time on dead ends, and better confidence in where to actually dig.

Luke

That's comparing it to one specific model—SwinV2-B trained on ImageNet. We don't know how it performs against other approaches, or whether that 22 percent improvement translates to real-world mission planning. The paper exists, but we're only seeing the headline number.

Mark

So scientists were really just looking at maps by hand before this?

Mimi

For the most part, yes. Or they'd use narrow models built for one specific task. Imagine having petabytes of data but no way to ask cross-cutting questions—no way to say, "Show me everywhere that looks like it might have ice, based on thermal signatures and topography combined." This model lets you ask those bigger questions.

Luke

The article says scientists had to "either sift through maps and images by hand or use low resolution, task specific machine learning models." That's a binary that might not be entirely fair. There were probably intermediate approaches. But the core point stands: there was no unified, multi-instrument framework like this before.

Mark

What does it mean that this is part of a "family" of foundation models?

Mimi

IBM and NASA are building a philosophy: instead of one-off tools for weather, or geospatial data, or the Moon, they're creating a shared infrastructure. A researcher working on climate might use Prithvi for weather data. A lunar scientist uses this model. They're all built on similar principles, so knowledge transfers between domains.

Luke

That's the vision they're articulating, but we don't yet know if it actually works that way in practice. The article describes the intention, not the outcome. We'll need to see whether researchers actually do transfer knowledge between these models or whether they remain separate tools.

  • Petabytes of lunar data sat largely untapped — scientists faced weeks of manual analysis or narrow, costly models that couldn't capture the Moon's full complexity.
  • The bottleneck threatened to slow planning for Moon bases and future Mars missions, where identifying water ice and safe landing sites is not optional but existential.
  • IBM and NASA responded by training a unified foundation model on over 30 data layers from nine instruments across four missions, creating the first publicly available AI tool built specifically for lunar science.
  • The model cuts error in lunar ice identification by 22%, improves volcanic feature mapping by 3%, and matches leading crater detection tools while using half the training data.
  • By open-sourcing both the model and its training dataset, the agencies are redistributing scientific capability — giving institutions of any size access to tools once reserved for the well-resourced.
  • The release lands as part of a growing IBM-NASA family of foundation models spanning weather, geospatial, and heliophysics data, signaling a structural shift in how scientific AI is built and shared.

For generations, humanity has gazed at the Moon and sent instruments to listen — accumulating a vast but largely inaccessible archive of lunar knowledge. Now, IBM and NASA have released an open-source AI foundation model trained on decades of multi-mission lunar data, offering scientists a shared language for translating raw observation into actionable understanding. The model improves detection of lunar ice deposits, craters, and volcanic features with measurable precision gains, and its open release invites researchers worldwide to build upon a common scientific infrastructure. It is a quiet but consequential step in the long arc from curiosity to sustained human presence beyond Earth.

For decades, NASA's instruments have been watching the Moon — mapping craters, measuring thermal signatures, scanning for ice in permanently shadowed regions. The data accumulated in petabytes across nine instruments on four separate missions. But having data and being able to use it are different things. Scientists either spent weeks manually sifting through maps and images or relied on narrow, task-specific models that were expensive and imprecise. The bottleneck was not observation; it was translation.

IBM and NASA have now released an open-source foundation model trained on this accumulated lunar record. The NASA-IBM Lunar Foundation Model is the first publicly available AI system built specifically for scientific exploration of the Moon, learning from a unified dataset of more than 30 spatially-aligned data layers drawn from NASA's Lunar Reconnaissance Orbiter, the GRAIL mission, and Japan's SELENE/Kaguya spacecraft. Where no common framework previously existed, the model offers one.

The practical results are concrete. In testing, the model reduced error in identifying high-potential lunar ice areas by 22% compared to existing state-of-the-art tools — a meaningful gain for missions where water and oxygen deposits are considered essential for sustaining human presence and producing rocket fuel for Mars. It captures volcanic surface features 3% more accurately than comparable models at lower fine-tuning cost, and for crater detection — critical for geology, landing site selection, and long-term infrastructure planning — it matches leading alternatives while outperforming them by nearly 19% at standard resolution using only half the training data.

NASA's chief science data officer Kevin Murphy described the release as a shift in how the agency approaches its archive: collecting data, he noted, is only part of the job. IBM Research's Juan Bernabe-Moreno emphasized the open platform dimension — the model connects observations across instruments, surfaces patterns difficult to see in isolation, and invites the global research community to build on the same foundation.

The lunar model extends a broader IBM-NASA collaboration that has already produced open foundation models for geospatial, weather, and heliophysics data. The underlying logic is the same: rather than building a new algorithmic system for each research question, scientists can now begin from a shared, pre-trained base and adapt it to their needs. By open-sourcing both the model and its training dataset, the agencies are democratizing access to tools once available only to well-resourced institutions — and accelerating the pace of discovery for everyone.

For decades, NASA's instruments have been watching the Moon—mapping its craters, measuring its thermal signatures, scanning for ice in permanently shadowed regions. The data accumulated in petabytes, layered across nine different instruments on four separate missions. But having the data and being able to use it are different things. Scientists either spent weeks manually sifting through maps and images or relied on narrow, task-specific machine learning models that were computationally expensive and often lacked the precision needed for serious scientific work. The bottleneck was not observation; it was translation.

IBM and NASA have now released an open-source foundation model trained on this accumulated lunar record, designed to help researchers surface patterns across decades of multi-instrument observations at a scale no single tool has previously offered. The NASA-IBM Lunar Foundation Model represents the first publicly available foundation model built specifically for scientific exploration of the Moon. It works by learning from a unified dataset that aggregates more than 30 spatially-aligned data layers from instruments aboard NASA's Lunar Reconnaissance Orbiter and GRAIL mission, supplemented by complementary observations from Japan's SELENE/Kaguya spacecraft. The result is a machine learning-ready dataset that brings together tens of thousands of images and maps showing different geophysical properties of the lunar surface and subsurface—a common framework where previously no such unified resource existed.

The practical applications are immediate and concrete. Researchers can use the model to identify potential ice deposits in the Moon's permanently shadowed regions, where water and oxygen lie buried beneath the surface. These resources are considered essential for establishing a sustained human presence on the Moon and for producing rocket fuel for future missions to Mars. In testing, the model reduced error in identifying areas with high potential for lunar ice by 22 percent compared to existing state-of-the-art models. The model also helps scientists study the Moon's volcanic history by identifying Irregular Mare Patches—volcanic features that reveal information about the Moon's thermal evolution and that matter strategically for future surface operations. It captures the extent of these features 3 percent more accurately than comparable models while requiring lower fine-tuning costs. For crater detection, which is essential for understanding the Moon's age, geology, and chemical composition, and for selecting safe landing sites and planning long-term infrastructure, the model achieves comparable accuracy to leading alternatives while offering greater efficiency. At context-scale resolution of roughly 100 meters, it outperforms existing models by nearly 19 percent using only half the training data.

Kevin Murphy, NASA's chief science data officer and acting chief data and AI officer, framed the release as a fundamental shift in how the agency approaches its scientific archive. "NASA has spent decades building an extraordinary scientific record of the Moon, but collecting data is only part of the job," he said. "We also have to make data easier for scientists to explore and use." Juan Bernabe-Moreno, director of IBM Research Europe, UK and Ireland, emphasized the collaborative dimension: "The NASA-IBM Lunar Foundation Model gives scientists a foundation to explore the Moon at scale, connecting observations across instruments, revealing patterns that are difficult to see in isolation, and providing an open platform the global research community can build on."

The release extends a broader partnership between IBM and NASA that has produced a family of open foundation models spanning geospatial data, weather, heliophysics, and now lunar science. The strategy reflects a shift in how scientific AI is being built: instead of constructing a new algorithmic system for each research question, scientists can now start from a shared, pre-trained model and adapt it to their specific tasks. By open-sourcing the lunar model, the agencies are inviting the global research community to build on the same foundation, accelerating discovery across domains and democratizing access to tools that were previously available only to well-resourced institutions. The model is available now, and the unified lunar dataset it was trained on is open to researchers worldwide.

Collecting data is only part of the job. We also have to make data easier for scientists to explore and use.
— Kevin Murphy, NASA Chief Science Data Officer
The model gives scientists a foundation to explore the Moon at scale, connecting observations across instruments and revealing patterns that are difficult to see in isolation.
— Juan Bernabe-Moreno, Director of IBM Research Europe
Vuoi la storia completa? Leggi l'originale su Portal ERP ↗
Contattaci Domande frequenti