In Rwanda, a study has asked one of the defining questions of our technological moment: can machines be trusted to judge the wisdom of other machines, especially when lives are at stake? The answer, drawn from 524 clinical queries evaluated by both human physicians and artificial intelligence, is that the machines are cheaper and more consistent — but they are not yet wise. They missed what the doctors caught, particularly the quiet distortions of demographic bias, reminding us that efficiency and judgment are not the same virtue.
AI judges fail to catch clinical bias; human experts still essential for medical AI oversight
Potential harm to patients in resource-constrained settings if biased clinical AI responses bypass human oversight due to cost pressures.