Health AI is getting squeezed by its own contradictions. Clinicians want explainable AI, but new research on clinician trust in AI finds that explanations make surprisingly little difference. Civil rights groups want proof that clinical AI models work equally well across race, geography, and income, but the data foundation behind most health AI still doesn’t include everyone it’s meant to serve.
Explainable AI in healthcare
Explainable AI, often shortened to XAI, refers to AI systems designed to show their reasoning rather than simply returning a prediction. In clinical settings, that usually means surfacing which factors, lab values, imaging features, patient history, drove a given output. The pitch has always been straightforward: if a black box model shows its work, clinicians will trust it more and use it more safely. That pitch is now being tested directly, and the results are more complicated than the pitch suggested.
The scrutiny isn’t coming from one direction. It’s coming from clinical researchers studying trust itself, and from civil rights organizations examining what’s actually inside the data these models are trained on. Both arrived at their conclusions independently. That’s what makes the pattern worth paying attention to.
The trust paradox clinicians are living with
Start with the clinicians, because they’re the ones being asked to use these tools every day. A recent study tested what actually happens when explanations get added to an AI model’s predictions, rather than assuming explainability improves clinical decision-making by default. Model predictions on their own reduced clinician error. Adding explanations produced a further improvement, but not a statistically significant one. The more telling finding was underneath that average: the impact of explanations varied sharply across participants, with some clinicians performing worse with explanations than without. Explanations increased clinician confidence. They had no significant effect on trust or reliance on the model itself [1].
That distinction matters more than it sounds. A clinician who feels more certain isn’t the same as a clinician who has good reason to be certain. Explainable AI, as an industry response to the black box problem in medicine, has mostly been sold as the second. This study is evidence it’s often only delivering the first. If that holds up, the entire premise of “clinicians don’t need black boxes, they need explanations” gets a lot shakier. Confidence without calibration is its own kind of risk in a clinical setting, and it’s exactly the kind of risk that’s hard to catch until something goes wrong.
Where the black box actually breaks: data representativeness, not model design
Which is exactly where the NAACP and Sanofi’s ACE Your Health AI Task Force is pointing [2]. Its recent health AI governance paper argues the deeper problem sits upstream of explanations entirely, in the data foundation the model was built on in the first place. Dermatology models trained on datasets where darker skin tones are a fraction of the images, despite being close to half the population, carry performance gaps of up to 50 percent on the conditions that matter most [2]. That’s not a one-off finding. Separate research reviewing clinical AI datasets more broadly found similarly stark numbers: Hispanic patients make up just 2.8 percent of training datasets, against 18 percent of the population. Black patients make up 7.3 percent, against 13 percent of the population [3].
The task force’s real contribution isn’t cataloguing the gap. It’s showing how the gap gets built in, stage by stage, across the entire AI development lifecycle. Using a Type 2 diabetes risk model as a worked example, the paper follows bias through five stages. It starts with how the question gets framed, then who ends up represented in the data, how the model is built and tested, whether clinicians are ever trained to catch the model when it’s wrong at deployment, and lastly, whether anyone is watching afterward for accuracy to quietly erode in exactly the populations the model was never built to serve. Their answer is structural rather than technical: mandatory bias audits, public model documentation, and subgroup performance data made available before deployment, not patched in after something goes wrong [2].
Put next to the clinician-trust study, the two pieces of research are making complementary arguments rather than competing ones. One says explanations don’t reliably produce trust. The other says the reason may be that there’s often nothing trustworthy underneath the explanation to begin with. A well-explained decision built on a thin, unrepresentative data foundation is still a fragile decision. It’s just a fragile decision that sounds more convincing, which arguably makes it riskier, not safer.
What health AI vendors need to prove before deployment
For anyone building or buying clinical AI right now, the practical takeaway converges on a short list. Subgroup performance needs to be measured and disclosed, not assumed. Data provenance needs to be documented well enough that an auditor or a procurement team can trace the training data to its source and see who it represents. Explainability features, where they exist, need to be evaluated for whether they actually change clinical outcomes, not just whether they make an interface feel more transparent. And governance needs to sit earlier in the development cycle than most organizations currently place it, at the problem-definition stage.
None of this is a call to abandon explainability features. It’s a call to stop treating them as the fix. An explanation layer on top of a narrow, unrepresentative dataset is still a narrow, unrepresentative model. It’s just one that’s easier to trust by mistake.
What this means for the industry building these systems
This is a rare thing in health AI right now: two unrelated sources landing on the same message without needing to be coordinated. It’s also the argument PAICON has been building around since day one. A model earns clinical trust the same way a diagnosis does, by being traceable back to the evidence that produced it. That’s what a data foundation is for. Not a black box with a better explanation bolted on, but an ecosystem where the data, its provenance, and its representativeness can all be checked by the people whose care depends on it.
Clinical researchers and civil rights advocates didn’t set out to make the same point. But they’ve landed on this much: the fix for a black box was never a better explanation. It was always a better foundation underneath it.
References
-
Nicolson A, Bradburn E, Gal Y, Papageorghiou AT, Noble JA. The human factor in explainable artificial intelligence: clinician variability in trust, reliance, and performance. NPJ Digit Med. 2025;8:658. doi:10.1038/s41746-025-02023-0
-
National Association for the Advancement of Colored People, Sanofi. Building a healthier future: designing AI for health equity [Internet]. Baltimore (MD): NAACP; 2025 Nov 17. Available from NAACP website.
-
Osonuga A, Osonuga AA, Fidelis SC, Osonuga GC, Juckes J, Olawade DB. Bridging the digital divide: artificial intelligence as a catalyst for health equity in primary care settings. Int J Med Inform. 2025;204:106051. doi:10.1016/j.ijmedinf.2025.106051
