Back to the episode map

Evergreen

How to Run a Bounded Workplace AI Pilot

Run a workplace AI pilot with one task, a preserved baseline, representative cases, access controls, human review, incident rules, stop conditions, and a dated decision.

Aug 4, 20267 min readBy Dalton Anderson

How to Run a Bounded Workplace AI Pilot

A bounded workplace AI pilot tests one defined task, with named users, approved data, a preserved baseline, a fixed evaluation period, accountable human review, incident rules, and a decision date. It is not a quiet rollout with the word “pilot” attached.

The boundary protects the organization and improves the evidence. When the task, sources, users, and success criteria stay stable, the team can learn whether the feature actually improves work.

flowchart TD
    A["One work task and accountable owner"] --> B["Baseline and acceptance standard"]
    B --> C["Approved users, data, tenant, plan, and controls"]
    C --> D["Representative cases and harmful-failure tests"]
    D --> E["Time-boxed live pilot"]
    E --> F["Review outcomes, access, incidents, time, and cost"]
    F --> G{"Predefined decision rule"}
    G --> H["Approve a narrow production use"]
    G --> I["Revise and retest"]
    G --> J["Stop and retain evidence"]

Choose a task that can produce evidence

The first pilot should have a recognizable start, a reviewable result, a manageable consequence, and enough repetition to compare outcomes.

Meeting-note review, retrieval of an internal policy, classification of routine requests, or a first draft from approved source material can fit. “Transform how the company works” cannot.

Name the task in operational language. Identify who performs it today, who approves the result, who consumes it, how often it occurs, and what failure would mean. Exclude adjacent uses.

If the team is testing meeting recap, do not silently add contract drafting, personnel decisions, customer messaging, and enterprise search halfway through the pilot. Each use may require a different source boundary and risk decision.

Appoint an owner who can stop the test

The pilot owner is accountable for scope, configuration, evidence, participant instructions, incidents, and the final recommendation. The owner needs authority to pause the feature when the test crosses its boundary.

Supporting reviewers may include security, privacy, legal, records, accessibility, labor, procurement, information owners, and domain experts. The right group depends on the task and jurisdiction.

NIST’s AI Risk Management Framework organizes voluntary risk-management activity around govern, map, measure, and manage. A small pilot does not need a theatrical governance program, but it does need named decisions across those functions.

Preserve the existing baseline

Observe the current process before the pilot begins. Record completed quality, total time, review time, material error, rework, support burden, and cost using the same definitions the pilot will use.

Do not compare an AI-assisted result with an imaginary perfect manual process. Do not compare the AI’s first draft with the fully reviewed human result.

The unit of comparison is completed work. Both paths end when the result reaches the same acceptance standard and enters the same system of record.

Keep enough baseline examples to represent ordinary and difficult work. A pilot built only from clean, common cases will not reveal how the feature behaves when the evidence is missing, contradictory, stale, or restricted.

Freeze the technical and data boundary

Record the exact product, plan, tenant, feature state, model if visible, user roles, administrative policies, connected sources, storage location, retention rule, and documentation date.

Microsoft distinguishes Microsoft 365 Copilot from Copilot Chat and describes different organizational-data behavior across the products in its current overview. Slack’s current AI guide ties capabilities to plans, roles, and administrator restrictions.

That configuration is part of the product being tested. A license change or newly connected source can invalidate the result.

Create an allowed-data rule. State which information may enter the feature, which information is prohibited, which output destinations are permitted, and how users should handle an accidental disclosure.

Repair obvious access problems first

An AI search or summary feature can expose the consequences of weak information architecture. It can retrieve an old policy, blend duplicate records, or surface a broadly shared document that should have been restricted.

Review the source corpus and its permissions before interpreting pilot results. Identify the authoritative records, stale material, duplicate files, inherited access, external sharing, and sensitive repositories.

Microsoft’s security and privacy documentation says Microsoft 365 Copilot works within existing user permissions. Slack’s AI security documentation similarly describes its current data controls and commitments.

Permission-aware retrieval is necessary. It does not determine whether the permission itself is correct.

Build the evaluation set before live use

Create representative cases with known outcomes. Include ordinary inputs, ambiguous requests, missing evidence, conflicting sources, stale sources, uncommon names, restricted content, prompt injection or manipulative source text where relevant, and a case where the system should refuse or ask for clarification.

For every consequential case, preserve the controlling source and acceptable answer. Define harmful failure in advance.

A meeting-note pilot might treat an invented decision, wrong owner, incorrect deadline, erased dissent, or sensitive disclosure as harmful. An internal-search pilot might treat a restricted-source leak, unsupported policy answer, or failure to identify conflicting versions as harmful.

The NIST Generative AI Profile provides a current reference for risks such as confabulation, information integrity, privacy, and human overreliance. It should inform the cases, not replace task-specific judgment.

Train participants on the boundary

Participants need a short instruction that explains the approved task, allowed data, required review, prohibited uses, feedback method, incident path, and stop condition.

Tell them that fluency is not evidence. Require them to open citations and compare consequential statements with the source. Make it safe to report a bad result without being treated as resistant to the technology.

Training should also explain what the system cannot see. A user who assumes the feature has project history, meeting chat, a shared screen, or the latest policy may trust an incomplete output.

Run the pilot long enough to find variance

A one-hour demonstration tests whether the interface can produce an output. A pilot tests whether people can use it across real variation.

Use a fixed period or a fixed number of cases. Track who used the feature, the work type, source conditions, result, review time, corrections, access behavior, incidents, user confusion, and downstream rework.

Do not reward participants for maximizing prompts or outputs. High usage can reflect novelty, repetition, or poor initial quality. Measure whether completed work improved.

Exercise the failure path

The pilot should deliberately test what happens when the feature is wrong or unavailable. Confirm how a user reports an incident, how access can be restricted, how generated material is contained, how downstream records are corrected, and how the team returns to the existing process.

Set immediate stop conditions. These can include exposure of restricted information, an uncontained harmful output, repeated failure to follow the source boundary, loss of required records, or use outside the approved task.

A stop is not a failed experiment. It is evidence that the current combination of product, configuration, workflow, and controls is not ready.

Make a dated decision

At the end, compare completed outcomes, harmful failures, total time, correction burden, support, cost, access behavior, and participant experience with the baseline.

The decision should be one of four states: approve the exact use with named controls, narrow and retest, pause pending a defined change, or reject the use.

An approval should state the users, task, data, tenant, plan, sources, review requirement, monitoring owner, incident process, and refresh trigger. It should not become blanket permission for every feature in the suite.

Products and organizational data both change. Set a review date and retest after a material model, plan, integration, policy, permission, source, or workflow change.

[[How to Evaluate a Workplace AI Feature]] contains the full measurement method. [[How to Test AI Search and Summaries Against Sources]] provides the evidence test for retrieval-heavy pilots.

The pilot is complete when the organization has a defensible decision, not when the trial license expires.

Editorial note

This guide was developed with AI assistance from the immutable E042 transcript and the linked NIST, Microsoft, Slack, pilot, security, privacy, and product records. Dalton Anderson remains the author. Security, privacy, legal, accessibility, labor, procurement, records, domain, source, and founder review are mandatory before publication. Publication is not authorized.

Sources

Follow the evidence.

  1. slack.com: 28244420881555 Manage access to AI features in Slackslack.com
  2. learn.microsoft.com: recording transcription overviewlearn.microsoft.com
  3. open.spotify.com: 0FyyANPnMYdcc04GiM2OWXopen.spotify.com
  4. daltonanderson.ghost.io: ai in the workplace is copilot and slack ai worth itdaltonanderson.ghost.io
  5. learn.microsoft.com: security microsoft 365 copilotlearn.microsoft.com
  6. iea.org: key questions on energy and aiiea.org
  7. iea.org: data centre electricity use surged in 2025 even with tightening bottlenecks driving a scramble for solutionsiea.org
  8. slack.com: 31377193680019 Use AI to take huddle notes in Slackslack.com
  9. NIST AI Risk Management Frameworknist.gov
  10. iea.org: executive summaryiea.org
  11. slack.com: 115004846068 Slack updates and changesslack.com
  12. learn.microsoft.com: microsoft 365 copilot overviewlearn.microsoft.com
  13. slack.com: 28310650165907 Security for AI features in Slackslack.com
  14. youtu.be: ZMvMBflUd4youtu.be
  15. NIST Generative AI Profilenvlpubs.nist.gov
  16. slack.com: 25076892548883 Guide to AI features in Slackslack.com
  17. daltonanderson.net: ai in the workplace is copilot and slack ai worth itdaltonanderson.net
How to Run a Bounded Workplace AI Pilot