AI model identifies lead contamination risk across cities with limited data

Lead contamination in residential water poses persistent public health risks, particularly affecting vulnerable populations in data-limited jurisdictions.
A model that works with messy real-world data and still identifies contamination hotspots
The significance of the research lies not just in accuracy but in its potential deployment in under-resourced cities.
Mark

Why does it matter that the model understands water systems as networks rather than treating each house separately?

Mimi

Because lead contamination doesn't respect property lines. If a service line is corroded, it affects everyone connected to it. A tabular model sees each house as independent data points. A graph model sees the actual plumbing—the relationships, the flow, the shared vulnerability.

Mark

So the model learned something about how cities are built?

Mimi

Exactly. It learned that proximity matters, that shared infrastructure creates shared risk. When it moved from Flint to Andover, it carried that understanding with it. It didn't need to relearn that water systems have structure.

Mark

The accuracy was 81 percent. That means one in five predictions could be wrong. Is that good enough?

Mimi

It depends on the use case. If you're a city with no testing data at all, 81 percent tells you where to send your inspectors first. It's a screening tool, not a diagnosis. It narrows the search space dramatically.

Mark

What about the communities that need this most—the ones with almost no data?

Mimi

That's the hard limit. The model needs some baseline to learn from. But most American cities have at least some testing history, even if it's incomplete. The real barrier is usually not data; it's resources to act on what the data shows.

Mark

Does this solve the lead problem?

Mimi

No. It makes the problem visible. That's the first step.

  • Lead contamination in residential water remains a persistent and largely unmapped threat, with thousands of municipalities lacking the data to know where the danger is greatest.
  • Traditional detection methods treat each property in isolation, missing the fundamental truth that contaminated infrastructure spreads risk across entire neighborhoods and blocks.
  • Researchers built a self-supervised graph attention network that models the spatial relationships between homes sharing water infrastructure, achieving 81% accuracy in identifying at-risk properties.
  • The model transferred successfully from Flint, Michigan to Andover, Massachusetts without retraining, suggesting it can cross jurisdictional boundaries and work with the incomplete data most cities actually have.
  • The approach is now positioned as a screening tool for under-resourced communities—not a perfect solution, but a structured way to prioritize where to look first before the next crisis becomes undeniable.

Beneath the streets of American cities, lead pipes carry both water and an invisible burden of risk—one that falls hardest on communities least equipped to find it. Researchers have now trained an artificial intelligence to read the hidden geometry of water infrastructure, learning from one city's crisis to illuminate risk in another. By treating water systems as what they truly are—networks of shared fate rather than isolated properties—the model achieves meaningful accuracy even where data is scarce. It is a quiet but consequential step toward making an old and stubborn danger visible before it harms another generation.

Lead sits in the pipes beneath American cities, invisible and patient, leaching into drinking water through corroded service lines and old fixtures. In places like Flint, Michigan and Andover, Massachusetts, public health officials face the same frustrating reality: spotty data from homeowners who bothered to sample their water, and no reliable way to predict which neighborhoods are at highest risk.

Researchers have now built a machine learning model that changes this equation. Rather than treating each property in isolation, they created a graph attention network—an AI that understands how homes connect through shared water infrastructure. A contaminated pipe doesn't affect just one house; it affects everyone downstream. The model, trained on data from nine Flint wards, learns the spatial logic of contamination spread.

The harder test was whether knowledge gained in Flint could help a completely different city. It could. Applied to Andover without additional training, the model achieved 81 percent accuracy and caught two-thirds of contaminated properties—performing nearly as well as approaches requiring far more labeled data than most cash-strapped municipalities possess. The spatial graph representation consistently outperformed traditional tabular models that ignored the networked nature of water systems.

What makes this significant is not the accuracy figure alone, but the pathway it opens. Flint and Andover are not anomalies—they are examples of a much larger problem facing thousands of American municipalities with aging infrastructure and incomplete testing histories. A model that can transfer knowledge across cities, work with messy real-world data, and still identify contamination hotspots offers a structured way forward. The spatial logic of water systems, the way pipes connect neighborhoods and contamination spreads, can be made legible to machines in ways that help humans see the risk before it becomes a crisis.

Lead sits in the pipes beneath American cities, invisible and patient. It leaches into drinking water through corroded service lines and old fixtures, and no one knows exactly where the problem is worst until someone tests the water—if they test it at all. In places like Flint, Michigan, where the water crisis became national news, and in smaller cities like Andover, Massachusetts, where contamination was discovered more quietly, public health officials face the same frustrating problem: they have spotty data from homeowners who bothered to sample their water, and no reliable way to predict which neighborhoods are at highest risk.

Researchers have now built a machine learning model that changes this equation. Instead of relying on scattered homeowner samples and traditional statistical tables that treat each property in isolation, they created a graph attention network—a type of artificial intelligence that understands how properties connect to one another through shared water infrastructure. The model learns by studying the spatial relationships between homes on the same block, the same street, the same water system. A contaminated pipe doesn't affect just one house; it affects everyone downstream.

The team trained their self-supervised graph attention network, or SSGAT, using data from nine different wards across Flint. They then tested whether what the model learned in Flint could work in Andover, a completely different city with its own water system, its own pipe network, its own contamination patterns. This is the hard test—can a model trained on one city's problem actually help another city?

It did. When applied to Andover without any additional training, the model achieved 81 percent accuracy in identifying which properties were at risk, with a macro-recall score of 0.66. That means it caught two-thirds of the contaminated properties, even when working with incomplete information. For comparison, the model performed nearly as well as a fully supervised approach that required much more labeled training data—the kind of data that most cash-strapped municipalities simply don't have. The spatial graph representation mattered. It outperformed traditional tabular models that ignored the fact that water systems are networks, not isolated points.

The researchers tested their predictions against multiple contamination thresholds, including the EPA's current action level of 10 parts per billion. At that threshold, the model's ability to flag at-risk properties held up. This matters because 10 ppb is the level at which the EPA says water systems must take action—must notify residents, must offer treatment, must investigate. But many cities don't know where to look first.

What makes this work significant is not just the accuracy number. It's the pathway it opens for cities with limited resources. Flint and Andover are not anomalies; they are examples of a much larger problem. Thousands of American municipalities have aging water infrastructure, incomplete testing data, and no systematic way to prioritize which neighborhoods need immediate attention. A model that can learn from one city's experience and transfer that knowledge to another, that can work with messy real-world data and still identify contamination hotspots with reasonable confidence—that is a tool that could actually be deployed.

The next question is whether it will be. The model works best when there is at least some baseline data to learn from, some testing history. In the most under-resourced communities, even that may not exist. But for the many cities caught between crisis and invisibility, between knowing there is a problem and knowing where it is, this approach offers a structured way forward. It suggests that the spatial logic of water systems themselves—the way pipes connect neighborhoods, the way contamination spreads—can be made legible to machines in ways that help humans see the risk more clearly.

Spatial graph representations can support cross-city generalization under class-imbalanced conditions, offering a structured approach for contamination screening in data-limited jurisdictions
— Research findings
Contact Us FAQ