Back to the episode map

Guide

How to Measure Workplace AI Productivity and ROI

Measure accepted work, total effort, quality, rework, harmful failures, adoption, worker impact, cost, and uncertainty before making a workplace AI ROI claim.

Aug 4, 20267 min readBy Dalton Anderson

How to Measure Workplace AI Productivity Without Inventing ROI

Measure workplace AI against accepted work, not generation speed. Compare comparable tasks under the same completion standard and count setup, source gathering, prompting, waiting, review, correction, rework, support, incidents, and cost. Report quality and harmful failures beside time.

A fast draft can create a slow process. A positive average can hide a damaging tail. A worker's estimate of time saved can be useful feedback without becoming a financial return.

flowchart TD
    A["Define accepted work"] --> B["Preserve the current baseline"]
    B --> C["Match representative tasks and conditions"]
    C --> D["Measure total effort, quality, and harmful failure"]
    D --> E["Add training, support, licenses, controls, and incidents"]
    E --> F["State uncertainty and worker-level variation"]
    F --> G{"Approve, modify, pause, or stop"}

Define the completed outcome

Choose one task and write the acceptance standard before the test.

For a meeting record, accepted work may require verified decisions, owners, deadlines, conditions, dissent, and a destination. For a customer response, it may require an authorized claim, correct account context, policy fit, privacy, tone, and approval. For a support ticket, it may require reproducible steps, evidence, severity, and routing.

The AI-assisted clock does not stop when text appears. It stops when the result reaches the same standard as the baseline and enters the intended system.

Do not compare a rough human draft with a polished AI output, or an AI first draft with a fully reviewed human record. Compare equivalent completed work.

Preserve the baseline before the workflow changes

Collect a representative sample from the existing process. Record active effort, elapsed time, accepted quality, material error, rework, review, support, and downstream correction.

Use ordinary and difficult cases. Include missing evidence, conflicting sources, unusual terminology, high-volume periods, handoff problems, and cases that should be escalated.

The baseline can be imperfect. If the current process is unreliable, preserve that fact. The decision is whether the assisted process improves the real starting point enough to justify its full cost and risk.

Document who performed the work, their experience, which sources they used, and any accessibility or workload conditions that affect interpretation. Averages across different people and tasks can blur the result.

Choose measures that follow the work

Use a compact scorecard.

MeasurePractical definition
Accepted outcome rateShare of cases meeting the full completion standard
Material correctnessConsequential claims verified against controlling evidence
CompletenessRequired elements retained, including uncertainty and dissent
TraceabilityConsequential claims connected to supporting sources
Active effortHuman minutes spent across the entire path
Elapsed timeTime from trigger to accepted record
Review and correctionReviewer minutes and material changes
ReworkWork repeated after initial acceptance or downstream discovery
Harmful failurePredefined error capable of changing a decision or harming a person
Adoption burdenTraining, help, workarounds, and support
Worker impactWorkload, autonomy, accessibility, stress, and role change
Total costLicense, integration, administration, controls, and incidents

Do not let one composite score erase tradeoffs. A process that is faster but less accurate is not automatically better. A process that improves output but creates inaccessible work or unrecorded surveillance needs a different decision.

Match the comparison

Use similar tasks, sources, difficulty, deadlines, and reviewers across the baseline and assisted conditions. Record the product, plan, model if visible, configuration, connected sources, account type, and date.

Random assignment can improve causal evidence when it is ethical and practical, but a workplace pilot may not have the sample or authority for a formal experiment. If the design is observational, say so. Do not imply randomization, blinding, or statistical power that did not exist.

Track non-use and dropout. If only enthusiastic volunteers use the product, the result may not represent the broader workforce. If the system changes during the test, record the change rather than blending versions.

Blind quality review can reduce some expectation effects when reviewers can judge the work without seeing the condition. It cannot remove every source of bias, and it may be inappropriate when the workflow needs disclosure.

Learn from research without importing its percentage

Workplace evidence is task-specific.

Brynjolfsson, Li, and Raymond's customer-support field study examined 5,179 agents and reported an average increase in issues resolved per hour. The effect differed by prior worker experience. The measure and system were specific to customer support.

Dillon and colleagues' field experiment across 66 firms reported that users of an integrated generative AI tool spent less time on email and less time working outside regular hours. The researchers did not detect a change in the quantity or composition of tasks from individual access alone. The paper also discloses relevant Microsoft relationships.

Dell'Acqua and colleagues' consulting experiment found different effects across tasks, including worse performance where the system was outside the tested capability boundary.

The lesson is not to choose the most attractive percentage. It is to test the exact task and watch for variation.

Count hidden review labor

Review can shift work rather than remove it.

Track who checks the output, how long verification takes, which sources they reopen, how many corrections they make, and whether the reviewer becomes a bottleneck. Record the emotional and cognitive burden of monitoring fluent but uncertain output.

If workers must use the tool while maintaining the prior process as a fallback, the pilot may temporarily increase workload. Preserve that cost instead of calling it resistance.

Material corrections should remain visible. Silent cleanup makes the system look more accurate than it was and prevents the team from seeing repeated failure patterns.

Treat failure severity separately from frequency

An average error rate can hide an unacceptable event.

Define harmful failure before the test. Examples can include a privacy exposure, invented customer commitment, discriminatory employment recommendation, wrong payment, unsafe instruction, false legal statement, unsupported public claim, or production defect.

Report the count, severity, detectability, containment, and downstream correction. A rare severe failure may outweigh many harmless gains.

The NIST Generative AI Profile can help identify risk categories. It does not decide the organization's tolerance.

Calculate cost without pretending every minute becomes cash

Total operating cost can include licensing, integration, procurement, administration, training, review, correction, support, data cleanup, access repair, monitoring, compliance, incidents, vendor management, and switching.

Time reduction is not automatically labor-cost reduction. Saved minutes may become capacity, lower overtime, faster response, better quality, or nothing at all. The financial treatment depends on demand, staffing, utilization, accounting, and what the organization actually changes.

State the benefit in the unit observed. If people completed accepted cases faster, say that. If they reported lower after-hours work, say that. If the company realized a verified cost reduction, document the calculation and assumptions.

Avoid annualizing a short pilot without accounting for seasonality, learning, support, model change, and scale effects.

End with a dated decision

Choose approve the exact use with controls, modify and retest, pause pending a named condition, or stop.

The decision record should state the task, users, data, system, configuration, baseline, sample, outcome, variation, harmful failures, full cost, worker feedback, uncertainty, owner, and refresh trigger.

E042's [[How to Evaluate a Workplace AI Feature]] provides the broader product evaluation. E050's [[How to Run a Public Goal Review Without Performance Theater]] is useful when public accountability pressures the team toward a flattering story.

The honest ROI result may be that a feature helps one task for one group under one configuration. That is enough to make a good decision.

This guide was developed with AI assistance from the preserved E019 transcript, the linked evidence record, and the cited NBER, Harvard, and NIST sources. Dalton Anderson remains the author. It is not financial, accounting, employment, legal, procurement, security, privacy, or research advice. Finance, research, risk, worker, accessibility, domain, source, and founder review are required before publication. Publication is not authorized.

Sources

Follow the evidence.

  1. NIST AI RMF Measure guidanceairc.nist.gov
  2. ftc.gov: ai companies uphold your privacy confidentiality commitmentsftc.gov
  3. youtu.be: 0cC1Ez33ryIyoutu.be
  4. daltonanderson.ghost.io: ai in the workplace a practical guide to get starteddaltonanderson.ghost.io
  5. NIST AI Risk Management Frameworknist.gov
  6. NIST AI Resource Centerairc.nist.gov
  7. eeoc.gov: prohibited employment policiespracticeseeoc.gov
  8. eeoc.gov: us eeoc and us department justice warn against disability discriminationeeoc.gov
  9. nber.org: w31161nber.org
  10. open.spotify.com: 7LIXDoSM2gG97vFGftskQsopen.spotify.com
  11. NIST Privacy Frameworknist.gov
  12. nber.org: w33795nber.org
  13. eeoc.gov: strategic enforcement plan fiscal years 2024 2028eeoc.gov
  14. NIST Generative AI Profilenvlpubs.nist.gov
  15. ftc.gov: start security guide businessftc.gov
  16. dol.gov: ten 07 25dol.gov
  17. hbs.edu: dell acqua et al 2026 navigating the jagged technological frontier 5c589c8c fbb5 458f b285 c944746cd717hbs.edu
  18. cisa.gov: cisa and uk ncsc unveil joint guidelines secure ai system developmentcisa.gov
How to Measure Workplace AI Productivity and ROI