Back to the episode map

Guide

How to Choose Your First Workplace AI Task Safely

Choose a first workplace AI task using permission, consequence, reversibility, reviewability, baseline, data, worker impact, and stop-rule gates.

Aug 4, 20266 min readBy Dalton Anderson

How to Choose a First Workplace AI Task

The best first workplace AI task is permitted, narrow, reversible, measurable, and easy for a qualified person to review. It has a known input, a visible output, a stable acceptance standard, and a stop rule. It does not directly decide someone's rights, employment, safety, medical care, money, or legal position.

Start with the task, not the chatbot.

flowchart TD
    A["Inventory repeated work"] --> B{"Permission and approved data?"}
    B -->|"No"| C["Remove from pilot"]
    B -->|"Yes"| D{"Bounded, reversible, and reviewable?"}
    D -->|"No"| C
    D -->|"Yes"| E["Score frequency, baseline, consequence, and reviewer capacity"]
    E --> F["Choose one task"]
    F --> G["Write acceptance and stop rules"]

Remove prohibited work first

Before scoring convenience, check authority. The employer's policy, the data owner, the contract, the system administrator, law, professional duties, and the worker's role can all limit use.

A task leaves the first-use list when the input is not approved for the exact system, the worker lacks authority, an error is difficult to detect, the result is hard to reverse, or no qualified reviewer can own the outcome.

High-consequence tasks generally require a different adoption path. Employment decisions, legal interpretation, medical decisions, safety controls, financial approvals, customer promises, security changes, regulated filings, and production code cannot become "low risk" because the prompt is short.

The NIST AI RMF Core asks organizations to document purpose, context, task, human oversight, expected costs and benefits, third-party components, and potential impacts. It is a voluntary framework, not a product approval, but it provides a useful first screen.

The EEOC's employment-practice guidance shows why assignments, evaluations, hiring, promotion, pay, discipline, and other employment decisions require more than a low-consequence experimentation label.

Write the task as completed work

"Use AI for emails" is too broad. "Create a first draft of the weekly internal project update from approved status notes for the project lead to verify" is closer.

Name the trigger, input, action, output, reviewer, destination, and definition of accepted work. Exclude adjacent uses.

An email draft for an internal update is not permission to generate performance feedback, customer commitments, legal notices, or public claims. A document summary is not permission to interpret a contract or policy.

The task should fit in one sentence without words such as transform, optimize, everything, or general assistant.

Check whether the task can be reversed

Reversibility means the team can detect a bad result, prevent or correct downstream harm, restore the prior process, and remove temporary access.

A private draft that remains outside the system of record is easier to reverse than an automated message sent to customers. Suggested issue labels are easier to undo than a closed case. A proposed meeting summary is easier to correct than a payroll, hiring, or safety decision.

Ask what happens after the output. If it triggers work, payment, notice, access, discipline, publication, code deployment, or a formal record, the consequence is no longer contained inside the interface.

Confirm that review is feasible

Human review is a capacity requirement, not a disclaimer.

The reviewer needs the source, domain knowledge, time, authority, and a decision rule. If the output is longer or more complex than the source, review may cost more than the task saves. If an error is persuasive and difficult to detect, a quick glance is not an adequate control.

Define the material checks before the test. For a project update, the reviewer might verify milestones, owners, dates, blockers, and customer commitments. For a code explanation, the reviewer might compare behavior with the repository and tests.

[[How to Review AI Output Before It Enters Real Work]] provides the acceptance record.

Choose work with a baseline

A useful first task has enough repetition to compare the existing and assisted paths. Preserve representative examples before the tool is introduced.

Measure total active effort, elapsed time, accepted quality, material errors, rework, reviewer time, and downstream correction. The baseline does not need to be elegant. It needs the same completion standard as the assisted process.

If the task happens once a year, has no accepted output, or changes every time, it may be valuable work, but it is a poor first measurement task.

Frequency is not automatically good. A frequent task with sensitive data or hard-to-detect error can multiply harm.

Score the remaining candidates

Use a small evidence table instead of intuition.

GateStrong first-task signalWarning
PermissionApproved task, system, account, and dataWorker is guessing about policy
BoundaryOne trigger, input, output, and destination"Help with anything"
ConsequenceDraft remains contained until approvalOutput acts on people or systems
ReversibilityPrior process can resume immediatelyDownstream effects are hard to undo
ReviewabilitySource and qualified reviewer existCorrectness is difficult to observe
BaselineComparable completed work existsOnly anecdotal time estimates exist
FrequencyEnough cases to find varianceOne polished demonstration
Worker impactBurden and accessibility are visibleUsage becomes hidden performance monitoring

Do not collapse the table into one magic score. Permission, data authority, reviewer capacity, and unacceptable consequence are gates. A high frequency score cannot compensate for failing them.

Define harmful failure

An ordinary defect might be awkward wording that the reviewer notices immediately. Harmful failure could be an invented commitment, leaked information, discriminatory recommendation, false source, wrong number, unsafe instruction, or incorrect record.

Write the harmful-failure definition before testing. State whether one occurrence stops the pilot, triggers escalation, or requires containment and correction.

The NIST Generative AI Profile describes several risk categories that can inform the test. The organization still has to translate them into the exact task and affected people.

Produce a one-page selection record

The record should name the task, current process, owner, allowed data, approved system, reviewer, acceptance standard, baseline, expected benefit, known risk, ordinary failure, harmful failure, test size, stop condition, and decision date.

If those fields cannot be filled, the task is not ready.

After selection, [[How to Map a Workflow Before Buying an AI Tool]] shows where the tool would enter the real process. For a fuller production evaluation, use E042's [[How to Evaluate a Workplace AI Feature]].

The goal of the first task is not to prove that the organization is innovative. It is to learn something reliable without creating an uncontrolled second problem.

This guide was developed with AI assistance from the preserved E019 transcript, the linked task-selection record, and current NIST sources. Dalton Anderson remains the author. It does not authorize a system, data use, employment practice, or high-consequence decision. Policy, privacy, security, accessibility, labor, legal, domain, records, and founder review may be required. Publication is not authorized.

Sources

Follow the evidence.

  1. NIST AI RMF Measure guidanceairc.nist.gov
  2. ftc.gov: ai companies uphold your privacy confidentiality commitmentsftc.gov
  3. youtu.be: 0cC1Ez33ryIyoutu.be
  4. daltonanderson.ghost.io: ai in the workplace a practical guide to get starteddaltonanderson.ghost.io
  5. NIST AI Risk Management Frameworknist.gov
  6. NIST AI Resource Centerairc.nist.gov
  7. eeoc.gov: prohibited employment policiespracticeseeoc.gov
  8. eeoc.gov: us eeoc and us department justice warn against disability discriminationeeoc.gov
  9. nber.org: w31161nber.org
  10. open.spotify.com: 7LIXDoSM2gG97vFGftskQsopen.spotify.com
  11. NIST Privacy Frameworknist.gov
  12. nber.org: w33795nber.org
  13. eeoc.gov: strategic enforcement plan fiscal years 2024 2028eeoc.gov
  14. NIST Generative AI Profilenvlpubs.nist.gov
  15. ftc.gov: start security guide businessftc.gov
  16. dol.gov: ten 07 25dol.gov
  17. hbs.edu: dell acqua et al 2026 navigating the jagged technological frontier 5c589c8c fbb5 458f b285 c944746cd717hbs.edu
  18. cisa.gov: cisa and uk ncsc unveil joint guidelines secure ai system developmentcisa.gov
How to Choose Your First Workplace AI Task Safely