Essay

Measuring data products by the decisions they enable

Engagement and comprehension miss the thing that matters, whether someone can act on what they saw. A framework for evaluating data products on decision quality.

Abstract

Most data products today are evaluated on engagement and comprehension. Did people look? Did they understand? Both those metrics miss the thing that ultimately matters, which is whether the person could act on what they saw. I built an evaluation platform for this exact purpose.11“From Data to Action” is a live experiment with 50+ professionals already enrolled; see the Lab section. This essay defines actionability and argues it belongs at the center of how we evaluate data products.

I. Engagement is a proxy that wandered

Engagement was always a stand-in. We measured dwell time and clicks because they were easy and correlated, loosely, with value. But a metric optimized long enough stops being a proxy and starts being the target,22A restatement of Goodhart’s law for product analytics. and a dashboard can be maximally engaging and still leave every decision exactly where it stood before the product was introduced.

Comprehension metrics are the more respectable cousin, and they fail more subtly. A user who aces a quiz about a chart has demonstrated recall, not readiness to act. The gap between "I understood the chart" and "I made a better call because of it" is where data products actually earn or waste their existence, and neither engagement nor comprehension can fully capture it.

The visualization research community has admitted this gap in its own literature. Dimara and Stasko's audit of the field found that decision-making tasks, supposedly the point of the whole enterprise, barely appear in the task taxonomies visualization research uses to evaluate itself.44Dimara and Stasko, “A Critical Reflection on Visualization Research: Where Do Decision Making Tasks Hide?” IEEE TVCG 28, no. 1 (2022): 1128–1138. We built a discipline around perception and comprehension and assumed decisions would follow. The assumption was never tested because the tests weren't designed for it.

II. Defining actionability

I define actionability as the change a data product produces in the quality of a downstream decision. It goes beyond whether the user felt informed, to whether the call they made was measurably better (more accurate, better calibrated, faster without loss of accuracy) than the call they would have made without using the product.

Measuring actionability in this sense requires two signals moving together, what users say they can do and what their recommendations are actually worth when scored independently. Confidence without correctness is the exact trap engagement metrics walk into.

This definition of actionability is deliberately behavioral. Whether the interface was delightful or the user felt empowered is beside the point. It asks one thing. Put the same person in front of the same decision with and without the product, and score the difference.

III. Measuring it

The experimental shape is theoretically simple (and deceptively hard to run well). Hold the data and the task constant, vary the interface, and measure the decision quality.33The 4 interface types tested span excel-style static tables, pudding-style scrollytelling narratives, interactive sandbox environments, and conversational AI experiences — all serve identical underlying data.

In my live experiment, 50+ professionals so far have worked the same analytical task through one of four interfaces. Each participant reports what they believe they can act on; their recommendations are then scored independently.

The conversational arm matters because it imports a known hazard. The human-AI decision literature keeps finding that assistance shifts confidence faster than it shifts accuracy, and that whether people engage critically with an AI's output is a strategic choice about effort rather than a fixed trait.55Vasconcelos et al., “Explanations Can Reduce Overreliance on AI Systems During Decision-Making,” PACM HCI 7, CSCW1 (2023). A chat interface over data is a decision-support system whether its makers say so or not, and it should be evaluated like one.

IV. What it changes

If you evaluate on decision quality instead of engagement, a lot of received wisdom inverts. The most “engaging” interface is often not the most actionable. Simplicity that would lose an engagement test can win a decision test. And the product goal moves from holding attention to spending it well.

Notes

  1. “From Data to Action,” live experiment (React, Supabase, Claude API, Python). See Lab.
  2. Charles Goodhart’s observation, popularized as “when a measure becomes a target, it ceases to be a good measure.”
  3. Four architectures, identical data and task, decision measured at the end.
  4. Evanthia Dimara and John Stasko, “A Critical Reflection on Visualization Research: Where Do Decision Making Tasks Hide?” IEEE Transactions on Visualization and Computer Graphics 28, no. 1 (2022): 1128–1138.
  5. Helena Vasconcelos et al., “Explanations Can Reduce Overreliance on AI Systems During Decision-Making,” PACM HCI 7, CSCW1 (2023).

Works Cited

Dimara, Evanthia, and John Stasko. “A Critical Reflection on Visualization Research: Where Do Decision Making Tasks Hide?” IEEE Transactions on Visualization and Computer Graphics 28, no. 1 (2022): 1128–1138.

Goodhart, Charles. “Problems of Monetary Management: The U.K. Experience.” In Monetary Theory and Practice. London: Macmillan, 1984.

Kahneman, Daniel. Thinking, Fast and Slow. New York: Farrar, Straus and Giroux, 2011.

Vasconcelos, Helena, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, and Ranjay Krishna. “Explanations Can Reduce Overreliance on AI Systems During Decision-Making.” Proceedings of the ACM on Human-Computer Interaction 7, CSCW1 (2023).

Let’s work together — 

I’m open to full-time roles starting January 2027 —
and I’m always free for a coffee.

apetronedesign@gmail.com © 2026 Crafted by Angela Petrone · Privacy-friendly, cookieless analytics ↑ Back to top