Machine learning cracks the code on self-healing concrete performance

Healing time dominated—the longer concrete sat, the more it healed.
Machine learning analysis revealed which factors most strongly influence whether cracks in self-healing concrete actually close.
Mark

So the machine learning model predicts how much a crack in concrete will heal. Why does that matter? Why not just test the concrete and see what happens?

Mimi

Because testing takes time and money. If you want to optimize a recipe—find the best mix of ingredients and conditions—you'd need to make and test hundreds of samples. A model that can predict outcomes lets you narrow down the promising combinations first.

Luke

But how good are these predictions, really? The best model got an R of 0.704 on test data. That means it's explaining about 70 percent of the variation. Thirty percent is still unaccounted for.

Mimi

True. But the researchers also found that the difference between their best model and the second-best wasn't statistically significant. So we're not talking about a breakthrough—more like a useful tool that works reasonably well.

Mark

What did the model actually learn about what makes concrete heal better?

Mimi

Healing time is the biggest factor—longer time means more healing, obviously. Then initial crack width—narrow cracks close more completely. The concrete mixture proportions matter too, especially the ratio of binder to aggregate.

Luke

And the model's confidence in those rankings? How stable are they across different subsets of the data?

Mimi

They ran ten-fold cross-validation and got an average R of 0.641, which shows variability. The model's performance depends on which data it trains on.

Mark

So if I'm an engineer trying to design better self-healing concrete, what do I actually do with this?

Mimi

You use it as a screening tool. Instead of guessing which formulations to test, you run them through the model first, see which ones predict good healing outcomes, then test those in the lab.

Luke

But you'd still need to test them. The model can't replace physical validation.

Mimi

Exactly. It's a data-driven shortcut, not a replacement for real-world testing.

  • Self-healing concrete holds enormous promise for aging infrastructure, but its behavior is governed by so many interacting variables that reliable prediction has remained out of reach.
  • Researchers assembled 1,271 measurements across six studies and pitted four machine learning approaches against one another to see which could best forecast crack-closure rates.
  • The most sophisticated model — Least Squares Boosting optimized with a grey-wolf-inspired algorithm — achieved the strongest results, yet statistical testing revealed its edge over simpler decision-tree models was not significant enough to declare a clear victor.
  • SHAP analysis cut through the model competition to reveal what actually matters: healing time is the dominant factor, followed by initial crack width and the binder-to-aggregate ratio in the concrete mix.
  • The models still leave meaningful variation unexplained, but they offer materials scientists a rapid, low-cost way to explore formulations before committing to expensive physical testing.
  • The work positions machine learning as a practical accelerant in the transition from experimental self-healing materials to real-world engineering solutions.

Concrete has always cracked, and engineers have long dreamed of a material that could mend itself before those fractures become failures. A team of researchers has now turned machine learning toward that dream, training models on thousands of measurements to predict how completely self-healing concrete will close its own wounds. Their findings — that healing time, crack width, and mixture proportions govern the process most powerfully — suggest that the path from laboratory curiosity to durable infrastructure may be navigable with the right computational tools.

Concrete cracks — it always does. The deeper question is whether the material itself could be designed to repair those fractures before they become structural failures. Self-healing concrete, embedded with agents that activate when cracks form, offers that possibility, but predicting how well it will actually work has remained stubbornly difficult. Too many variables shift the outcome at once: crack width, mixture recipe, healing agent quantity, elapsed time. A research team recently attacked this prediction problem by training machine learning models on data from thousands of samples, searching for patterns that human intuition tends to miss.

The team drew on 1,271 measurements from six separate studies, feeding models five key inputs — initial crack width, concrete mixture proportions, healing agent quantity, and elapsed healing time — to predict a single crucial output: what percentage of the crack actually closed. They tested four approaches, from simple linear regression to Least Squares Boosting optimized by an algorithm modeled on grey wolf hunting behavior. Linear regression explained only about 57 percent of the variation in outcomes. Neural network and decision tree models reached roughly 68 percent. The boosting model pulled ahead with an R value of 0.704 on test data and 0.810 on training data. Yet when the team ran rigorous statistical comparisons across ten cross-validation folds, the performance gap between the boosting model and the decision tree model proved not to be statistically significant.

What the models learned mattered more than which one won. Using SHAP analysis, the researchers found that healing time was the dominant driver of crack closure — the longer concrete sat, the more it healed. Initial crack width ranked second, with narrower cracks closing more completely. The binder-to-aggregate ratio came third, followed by fly ash content and, with smaller influence, water-to-cement and coarse aggregate ratios. The variables also interacted: the effect of one input often depended on the values of the others.

The practical implication is that machine learning can serve as a rapid exploration tool for materials scientists, allowing them to test combinations of ingredients and conditions computationally rather than through endless physical experiments. The models are imperfect and their predictions shift depending on which data they train on, but they offer a meaningful starting point for understanding a material system that defies simple rules. As infrastructure ages and the demand for durability grows, that predictive capacity could meaningfully shorten the road from laboratory experiment to engineered solution.

Concrete cracks. It always does. The question engineers have wrestled with for decades is whether the material itself could be taught to repair those fractures before they become structural failures. Self-healing concrete—embedded with agents that activate when cracks form—offers that possibility. But predicting how well it will actually work remains stubbornly difficult. The healing process depends on too many variables at once: the width of the initial crack, the exact recipe of the concrete mixture, how much healing agent was added, and how much time has passed. Change one ingredient and the outcome shifts. Researchers recently tackled this prediction problem by training machine learning models on data from thousands of concrete samples, hoping to find patterns humans might miss.

The team gathered 1,271 measurements from six separate studies, each one documenting a concrete sample's journey from cracked to healed. They fed the models five key inputs: the width of the crack when it first formed, the proportions of the concrete mixture, the amount of healing agent present, and the elapsed healing time. The output they wanted to predict was simple but crucial—what percentage of the crack actually closed. They tested four different modeling approaches, ranging from straightforward linear regression to more sophisticated techniques like artificial neural networks and a method called Least Squares Boosting, each optimized using an algorithm inspired by the hunting behavior of grey wolves.

The results showed clear winners and losers. Linear regression, the simplest approach, managed an R value of 0.572 on the test set—meaning it explained about 57 percent of the variation in healing outcomes. The neural network and decision tree models performed better, both reaching around 0.68. But Least Squares Boosting, the most complex of the bunch, pulled ahead with an R value of 0.704 on the test set. On the training data it had seen before, it reached 0.810. On completely new validation data, it held steady at 0.666. The gap between these approaches mattered, but the researchers wanted to know if it mattered enough. They ran a statistical test—a Wilcoxon signed-rank test across ten separate cross-validation folds—and found something surprising: the difference between the best-performing boosting model and the decision tree model was not statistically significant. The boosting model did edge out the others on one particular metric, but even that advantage vanished after accounting for multiple comparisons.

What the models actually learned proved more interesting than which one won. Using a technique called SHAP analysis, the researchers identified which inputs had the strongest influence on healing outcomes. Healing time dominated—the longer the concrete sat, the more it healed, which makes intuitive sense. Initial crack width came in second; narrower cracks closed more completely than wide ones. The ratio of binder to aggregate in the mixture ranked third. The amount of fly ash in the concrete showed moderate influence. The ratios of coarse aggregate to cement and water to cement had smaller effects overall. The researchers also looked at how these factors interacted with each other, finding that the relationship between one variable and healing outcome often depended on the values of others.

The work suggests that machine learning can serve as a practical tool for materials scientists trying to optimize self-healing concrete formulations. Rather than running endless physical tests, researchers could use these models to rapidly explore which combinations of ingredients and conditions are likely to produce the best healing performance. The models are not perfect—they still leave significant variation unexplained, and their predictions vary depending on which subset of data they train on. But they offer a starting point for understanding a complex material system that defies simple rules. As concrete infrastructure ages and the demand for more durable materials grows, the ability to predict and optimize self-healing performance could accelerate the transition from laboratory curiosity to practical engineering solution.

Machine learning, especially LSBoost + GWO, can serve as a data-driven exploratory tool for estimating crack-healing percentage and investigating the relative importance and interactions of factors associated with the healing behavior of self-healing concrete.
— Research findings
Möchten Sie die ganze Geschichte? Das Original lesen bei Nature ↗
Kontakt FAQ