Transparency is the product: what clinical AI gets wrong about expert users
Surfacing every model detail backfires with specialists. What earns trust is the right information at the right moment in the workflow rather than more of it.
Abstract
Clinical AI teams treat transparency as an amount (more exposed model internals, more confidence scores, more explanation) and assume that adding it builds trust. Working with pathologists on a diagnostic tool, I found the opposite. Dumping model detail on expert users reduced trust, because it added cognitive load without answering the question the expert was actually asking.11Drawn from a validation engagement at UChicago Medicine’s SlideFlow Labs; see the case study on this site. This essay argues that transparency is a product you design rather than a quantity you add.
I. The transparency reflex
When a team ships an AI system into a high-stakes setting, the instinct is to make it maximally legible. Surface the features. Show the probability. Publish the confusion matrix. The reasoning is intuitive. Experts are skeptical, so give them everything and let them judge.22On explanation as a human, contrastive act, see Miller (2019).
But transparency defined as exposure assumes the bottleneck is information supply. For expert users it rarely is. A thyroid pathologist doesn't lack information; they're drowning in it. What they lack is a fast, defensible way to decide whether this particular output belongs in this particular case.
The recent human-AI interaction literature has caught up with this. Vasconcelos and colleagues showed that whether an explanation helps at all is a cost-benefit calculation the user makes in the moment. People engage with explanations when engaging is cheaper than the risk of being wrong, and ignore them otherwise.44Vasconcelos et al., “Explanations Can Reduce Overreliance on AI Systems During Decision-Making,” PACM HCI 7, CSCW1 (2023). Explanation is an offer rather than a substance you pour on a user, and experts decline expensive offers.
We watched this play out in validation sessions. The first prototype followed the reflex faithfully, with model internals, confidence values, and technical vocabulary. Pathologists' trust ratings landed anywhere from 1 to 3, and some declined to rate the tool at all. Nothing about the model had failed. The disclosure had.
II. What experts do with information
Expertise is compression. A specialist has already internalized which signals matter and which are noise, so their scarce resource is attention rather than data. Hand them a panel of twelve model internals and you haven't empowered them. You've handed them a second diagnostic problem on top of the first.
In sessions, the useful moments were never the raw numbers. They were the moments where the tool surfaced one piece of evidence, at the point in the read where a decision was being made, in the vocabulary the pathologist already used.33Bussone, Stumpf, and O’Sullivan (2015): more explanation can erode trust depending on framing and timing. That is still transparency, just aimed at a decision.
This tracks with what the explanation literature has said for years. An explanation is a social act, contrastive and selective. A data dump is none of those things. People don't want the full causal chain. They want the answer to "why this, rather than what I expected." Design for that question and a single well-placed piece of evidence outperforms a dashboard of internals.
III. Trust is a workflow property
We talk about trust as a static attitude. A user either trusts a system or doesn't. In practice trust is produced moment to moment, inside a workflow, by whether the system shows up as helpful exactly when it's consulted and stays out of the way otherwise.
That reframing changes the unit of design from the model to the encounter. The question is no longer “how do we explain the model” but “what does the expert need to see at the instant they would otherwise hesitate.” Answer that, and trust ratings move. We watched them move. After the redesign, every survey item scored a 4 or 5, from a starting point where nothing scored above 3.
It also explains why trust surveys taken outside the workflow mislead. Ask a specialist whether they trust an AI tool in the abstract and you'll get their prior about AI. Ask them after they've worked a real case with it, at the moment the tool either helped or got in the way, and you get something closer to the truth. We timed every rating to the end of contextual sessions for exactly this reason.
The field has started calling the real goal appropriate reliance, following the AI when it's right and overriding it when it's wrong.55Schemmer et al., “Appropriate Reliance on AI Advice,” IUI (2023). Maximal trust is not the target. A pathologist who trusts a diagnostic model unconditionally is as dangerous as one who refuses to look at it. Transparency's actual job is calibration, and calibration happens per case and per moment rather than per product.
IV. Designing the glass box
The redesign replaced a black-box output with a glass box, the model’s reasoning surfaced at the decision point, in validated terminology, with two distinct modes for the pathologist and the technician whose workflows actually differ. None of it adds transparency in the quantity sense. It treats transparency as a designed product with a job to do.
The lesson generalizes past medicine. Any time you put a model in front of an expert, the failure mode is the same, confusing legibility with usefulness. The fix is the same too. Design the information rather than the disclosure.
Notes
- Engagement at UChicago Medicine (SlideFlow Labs), thyroid pathology validation. See the linked case study.
- Tim Miller, “Explanation in Artificial Intelligence: Insights from the Social Sciences,” Artificial Intelligence 267 (2019): 1–38.
- Adrian Bussone, Simone Stumpf, and Dympna O’Sullivan, “The Role of Explanations on Trust and Reliance in Clinical Decision Support Systems,” ICHI (2015).
- Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, and Ranjay Krishna, “Explanations Can Reduce Overreliance on AI Systems During Decision-Making,” Proceedings of the ACM on Human-Computer Interaction 7, CSCW1, Article 129 (2023).
- Max Schemmer, Niklas Kühl, Carina Benz, Andrea Bartos, and Gerhard Satzger, “Appropriate Reliance on AI Advice: Conceptualization and the Effect of Explanations,” IUI (2023).
Works Cited
Bussone, Adrian, Simone Stumpf, and Dympna O’Sullivan. “The Role of Explanations on Trust and Reliance in Clinical Decision Support Systems.” In 2015 IEEE International Conference on Healthcare Informatics, 160–169. IEEE, 2015.
Miller, Tim. “Explanation in Artificial Intelligence: Insights from the Social Sciences.” Artificial Intelligence 267 (2019): 1–38.
Schemmer, Max, Niklas Kühl, Carina Benz, Andrea Bartos, and Gerhard Satzger. “Appropriate Reliance on AI Advice: Conceptualization and the Effect of Explanations.” In Proceedings of IUI 2023. ACM, 2023.
Vasconcelos, Helena, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, and Ranjay Krishna. “Explanations Can Reduce Overreliance on AI Systems During Decision-Making.” Proceedings of the ACM on Human-Computer Interaction 7, CSCW1 (2023).
Amershi, Saleema, et al. “Guidelines for Human–AI Interaction.” In Proceedings of CHI 2019. ACM, 2019.