Clinical AI

Earning clinical trust for an AI diagnostic tool

HEALTHCARE AICLINICAL VALIDATIONRECOMMENDATIONS PRODUCTIONIZED

An AI diagnostic tool pathologists didn’t trust became one they rated 4 or 5 out of 5 on every trust dimension. On the strength of the findings, the client productionized the recommendations and expanded their model offering. They also revised their go-to-market.

The problem

SlideFlow Labs is a clinical AI team within UChicago Medicine. It builds diagnostic software that helps pathologists identify computational biomarkers, removing the need for costly external testing. Their first model ready for clinical validation distinguished PTC from NIFTP in thyroid pathology, a difference that matters for both diagnosis and prognosis. The product existed but had never been tested outside the team that built it.

An unvalidated AI diagnostic tool in a clinical setting is a liability. The client needed to know whether pathologists would actually adopt the product before investing further. I led the research and product validation effort and managed 2 designers. I worked alongside ML engineers and physician scientists, and coordinated sessions with 12+ pathologists across 3 hospitals.

The key question

Will pathologists trust, accept, and use this tool?

Constraints

TimelinePathologist sessions were scarce and non-recoverable. Every interaction had to be designed to extract maximum insight.
Stakeholders12+ thyroid pathologists across 3 hospitals, among the hardest clinical users to access.
Technical / domainFindings had to survive clinical review. Every claim needed to hold up to a physician scientist.
Business needAn unvalidated clinical AI tool is a liability. The client needed proof of trust before further development.
Pathologist working through the prototype at a microscopy workstation during a usability session
RESEARCH SESSION WITH A PATHOLOGIST

Key decisions

With pathologist access this scarce, the wrong research approach was expensive in a way the schedule could not absorb. Five decisions shaped what we tested and, in turn, what shipped.

Decision #1: Research despite scarcity

The temptation was to minimize user research given how hard pathologist time was to secure. We committed to it and designed every session to extract maximum insight per interaction, because a validation effort without the actual users is not validation.

Decision #2: One use case only

Contextual inquiry showed that what “trust” means for an AI tool depends entirely on the use case. We scoped to thyroid pathology only. Generalizing would have made every finding weaker.

Decision #3: Trust over features

Rather than asking “does this feature work,” we asked “does this tool earn the right to be in a pathologist’s workflow.” That reframe determined what we tested, and it is why the result was a glass-box interface instead of a better black box.

Decision #4: Challenge the problem

We investigated whether the product was solving the right problem at all, and found that some pathologists had more pressing workflow concerns the tool could address more immediately than the diagnostic use case it was built for.

Decision #5: Test the transparency assumption

The instinct was to surface everything the model knew. When we tested that assumption, it failed. The flood of information confused experts instead of reassuring them. What pathologists needed was the right information at the right moment in their workflow.

The solution

The product moved from a black-box output to a glass-box interface. Instead of a bare result, pathologists see the model’s reasoning surfaced at the point in their workflow where it matters, in terminology validated through usability testing rather than raw technical specs. The design separated two distinct user modes, pathologist and technician, based on actual workflow mapping instead of assumed roles.

The research also surfaced digital readiness as an adoption variable that differed by hospital type. And it gave the client positioning intelligence they did not have before: how pathologists actually perceive the product next to larger competitors.

Final prototype: test analysis view with attention heatmap and high-attention tiles surfacing the model's evidence
GLASS-BOX INTERFACE · FINAL PROTOTYPE

Outcomes

3 → 5
trust went from no survey item above 3, with some pathologists declining to rate at all, to every item at 4 or 5 after the redesign (fresh raters each round, all three hospitals)
→ PROD
recommendations productionized; model offering expanded; go-to-market revised

Pathologists rated the tool on trust, clarity, and credibility, and only after hands-on sessions. Each round used fresh raters from all three hospitals. The client productionized our recommendations and built specific elements of the glass-box interface into the product. They expanded their model offering to more use cases to position as a full-stack platform. And after learning that competitors were already promising their original value proposition, they revised their go-to-market.

What I would do next

Let’s work together — 

I’m open to full-time roles starting January 2027 —
and I’m always free for a coffee.

apetronedesign@gmail.com © 2026 Crafted by Angela Petrone · Privacy-friendly, cookieless analytics ↑ Back to top