For as long as hospitals have been asked to trust artificial intelligence with patient care, a quiet dilemma has persisted: genuine validation requires real patient data, yet real patient data cannot be freely shared. Google Cloud and MLCommons have now built a third way — a platform called MedPerf, running on confidential computing infrastructure, where medical AI models can be tested against authentic clinical datasets without any party, including the cloud provider itself, ever seeing what flows through the evaluation. It is an attempt to reconcile the competing obligations of privacy, scie
Google Cloud and MLCommons enable private medical AI testing without exposing patient data
A model that works well in one setting can fail in another
Why has testing medical AI on real patient data been so fraught until now?
Hospitals sit on datasets they can't easily share—privacy law, competitive concerns, liability fears. But developers need that variety to know if their models actually work across different populations and equipment. It's been a standoff.
And confidential computing breaks the standoff how?
By making the data invisible to everyone during testing. The model runs in an encrypted, isolated hardware environment. Neither the hospital, the developer, nor Google can see what's happening inside. Only the results come out.
But doesn't that make it harder to debug if something goes wrong?
That's the trade-off. You lose transparency for privacy. But the brain tumor research shows it's worth it—they found one model was 95 percent accurate at one hospital and 63 percent at another. That kind of variation is exactly what you need to catch before patients see the system.
So this is really about trust. Hospitals trust the process because their data stays private.
Exactly. And regulators can trust the benchmarks because they're based on real clinical data, not lab conditions. It's not perfect, but it's the first time both sides get what they actually need.
What happens if the system itself is compromised?
That's the bet Google and MLCommons are making—that the hardware isolation is strong enough. They're using Intel and NVIDIA confidential computing together. It's not foolproof, but it's the best option available right now.
Der Puls
- Medical AI developers have long been forced to choose between meaningful validation and patient privacy — a compromise that has quietly undermined confidence in clinical AI systems.
- MedPerf deploys hardware-isolated Trusted Execution Environments so that neither hospitals, developers, nor Google can access the sensitive data or proprietary model code being evaluated.
- Brain tumor research has already exposed a troubling gap: one AI model scored 95% accuracy at one institution and collapsed to 63% at another, revealing how dangerously misleading narrow benchmarks can be.
- The federated approach allows institutions to contribute rare patient scans — such as glioblastoma MRIs — without pooling them in a single vulnerable location, preserving data sovereignty while enabling broader testing.
- Regulators and clinicians now have a potential pathway to validation evidence grounded in real-world clinical conditions rather than controlled laboratory datasets, which could meaningfully accelerate safe AI adoption.
For as long as hospitals have been asked to trust artificial intelligence with patient care, a quiet dilemma has persisted: genuine validation requires real patient data, yet real patient data cannot be freely shared. Google Cloud and MLCommons have now built a third way — a platform called MedPerf, running on confidential computing infrastructure, where medical AI models can be tested against authentic clinical datasets without any party, including the cloud provider itself, ever seeing what flows through the evaluation. It is an attempt to reconcile the competing obligations of privacy, scientific rigour, and clinical trust at a moment when the stakes of getting medical AI wrong are becoming impossible to ignore.
For years, the developers of hospital AI faced a dilemma with no clean resolution: test on real patient data and risk exposing sensitive records, or validate on narrow datasets that may not reflect how a system actually performs in a clinic. Google Cloud and MLCommons have now opened a third path.
MedPerf — an open-source benchmarking platform created by MLCommons in 2023 — now runs on Google Cloud's confidential computing infrastructure, using hardware-isolated Trusted Execution Environments, encrypted memory, and hardened operating systems. The result is an environment where AI models can be evaluated against genuine patient data without hospitals, developers, or even Google seeing what passes through the system. The platform runs on A3 machines equipped with NVIDIA H100 GPUs, capable of handling the intensive demands of real clinical evaluation.
The platform is already demonstrating its value in brain tumor research. The Federated Tumor Segmentation initiative — led by researchers at Indiana University, Northwestern University, and the University of Alberta — has been validating AI models trained to identify tumors in brain MRI scans. Because conditions like glioblastoma are rare, no single hospital accumulates enough cases to test a model thoroughly. The federated approach lets institutions contribute their private scans without ever pooling them in one vulnerable location.
What the researchers found was striking: a model that achieved 95 percent accuracy at one institution fell to just 63 percent at another. The gap is a reminder that medical AI systems can fail silently when moved between institutions with different patient populations, scanning equipment, or clinical protocols — and that without real-world validation, neither hospitals nor regulators can know whether benchmark results mean anything at all.
MedPerf on confidential computing addresses two long-standing obstacles simultaneously — privacy and representativeness — and in doing so may shorten the distance between research and clinical deployment, provided institutions extend their trust and regulators accept the evidence it produces.
For years, the people building artificial intelligence for hospitals have faced an impossible choice: test your model on real patient data and risk exposing sensitive medical records, or keep the data locked away and validate your work on narrow datasets that may not reflect how the system will actually perform in a clinic.
Google Cloud and MLCommons have now opened a third path. They've launched MedPerf—an open-source benchmarking platform—on Google Cloud's confidential computing infrastructure, creating a space where medical AI models can be evaluated against genuine patient data without anyone, including Google, ever seeing either the data or the model code. The system uses hardware-isolated environments called Trusted Execution Environments, encrypted memory, and hardened operating systems to ensure that hospitals, researchers, developers, and cloud providers remain blind to the sensitive information flowing through the evaluation.
The problem this solves is both practical and profound. Hospitals and research institutions operate under strict privacy regulations and understandably want to protect their own datasets. Developers, meanwhile, need access to varied patient populations to know whether their models work reliably across different institutions, different scanning equipment, different demographics. MedPerf, which MLCommons created in 2023 as part of a broader effort to standardize medical AI assessment, now runs on Google Cloud A3 machines equipped with NVIDIA H100 GPUs—hardware powerful enough to handle the intensive computational demands of testing AI models on real clinical data.
The platform is already proving its worth in brain tumor research. The Federated Tumor Segmentation initiative, led by researchers including Dr. Spyridon Bakas at Indiana University, Dr. Yury Velichko at Northwestern University, and Dr. Amber Simpson at the University of Alberta, has been using the system to validate AI models trained to identify tumors in brain MRI scans. Brain tumors like glioblastomas are rare enough that no single hospital accumulates enough cases to thoroughly test a model. The federated approach lets institutions contribute their private scans without pooling them in one vulnerable location.
What the researchers discovered was striking: a model that achieved 95 percent accuracy at one institution performed at only 63 percent accuracy at another. The gap reveals something crucial about medical AI—a system that works well in one setting can fail in another due to differences in patient populations, scanning protocols, or equipment. This variation is exactly why testing on representative, real-world data matters before a model enters clinical use. Without it, hospitals and regulators have no reliable way to know whether benchmark results reflect actual clinical performance or merely laboratory conditions.
Dr. Velichko described the experience of testing federated learning workflows on the new infrastructure as a shift from controlled laboratory settings to production-ready environments. Alexandros Karargyris, MedPerf's lead at MLCommons, framed the deployment as a step toward trustworthy validation—a way for clinicians, researchers, and regulators to have confidence in how medical AI is being assessed. The promise is substantial: hospitals and developers can now gather independent evidence of model performance without surrendering datasets or proprietary code, and regulators can review benchmarks grounded in real clinical conditions rather than narrow test environments.
The infrastructure represents a convergence of two technical challenges that have long complicated medical AI development. One is privacy—how to evaluate models without exposing patient data. The other is representativeness—how to test on data diverse enough to predict real-world performance. By solving both simultaneously, MedPerf on confidential computing may accelerate the path from research to clinical deployment, provided institutions trust the system and regulators accept the validation it produces.
Bemerkenswerte Zitate
The future of medical AI lies in secure, scalable, and collaborative cloud environments that move beyond controlled lab settings to test workflows in production-ready infrastructure.— Dr. Yury Velichko, Associate Professor of Radiology at Northwestern University
Medical AI holds enormous promise, but that promise can only be realized if clinicians, researchers, and regulators can trust the benchmarks we use to evaluate it.— Alexandros Karargyris, MedPerf Lead at MLCommons