Article
Start AI Adoption With a Bounded, Reviewable Task
Choose a first AI use case by defining the task, data, baseline, acceptable failure, reviewer, evidence, measures, owner, rollback, and exit.
Start AI Adoption With Bounded Tasks and Review
The best first AI use case is a bounded, reversible task with known inputs, a measurable baseline, permitted data, an explicit reviewer, visible evidence, unacceptable-failure tests, an accountable owner, and a clear exit. A broad instruction to "use AI" creates activity. A well-designed task creates evidence.
This is a method for selecting and running a pilot. It is not authorization to put confidential, personal, regulated, customer, employment, or proprietary information into a tool. Applicable policy, contracts, law, security requirements, and product terms still control.
flowchart TD
A["Define one task and baseline"] --> B["Set data and authority boundaries"]
B --> C["Name unacceptable failures"]
C --> D["Design representative tests"]
D --> E["Assign qualified review"]
E --> F["Measure quality, effort, risk, and recovery"]
F --> G{"Evidence supports expansion?"}
G -->|Yes| H["Expand one boundary"]
G -->|No| I["Revise, pause, or exit"]
1. Choose a task, not a department
Start with one unit of work that a person can describe from beginning to end. Drafting a response from an approved knowledge base is a task. Improving customer service is not. Extracting named fields from a known document type is a task. Automating operations is not.
Write the input, expected output, user, point of use, excluded actions, and current owner. State whether the system may suggest, draft, classify, retrieve, stage, or execute. Those verbs carry different authority.
A narrow task makes failure observable. It also limits how much workflow, data, and organizational change must happen at once.
2. Record the current baseline
Measure the existing task before evaluating the new one. Record completion time, wait time, error and rework rates, reviewer effort, escalation volume, cost, user experience, and any quality measure that matters.
The baseline does not need to be perfect. It needs to be honest enough to show whether the pilot improved the work or merely moved effort to another person.
AI output can appear fast while creating extra verification, correction, documentation, and recovery work. Compare the whole task, not only the moment when text appears.
3. Define the data boundary
List the information the pilot may use, the source of record, permitted purpose, retention expectation, access conditions, and prohibited data.
Check the provider's current product terms, security documentation, privacy commitments, data-use settings, regional availability, and administrative controls. A consumer account and an organization-managed service may have different protections. A public model announcement does not settle those questions.
The NIST Generative AI Profile treats risk management as a lifecycle activity. Its scope includes design, development, use, and evaluation rather than only model selection.
4. Name unacceptable failures before testing
Ordinary quality measures are not enough. Write the failures that would make the use unsafe, unlawful, misleading, inaccessible, economically unsound, or operationally unacceptable.
The list will depend on the task. It may include fabricated evidence, disclosure of protected information, discriminatory treatment, unauthorized commitments, incorrect customer instructions, missing required language, unsafe code, or an action taken without approval.
Create tests for those failures before a persuasive demonstration changes the team's risk tolerance.
5. Build a representative evaluation
Test realistic inputs, including ordinary cases, difficult cases, ambiguity, missing information, conflicting sources, uncommon formats, and known edge conditions.
Preserve the prompt or instruction, system and model version when available, source material, output, reviewer decision, correction, duration, and date. Without that record, a later team may not know what was actually evaluated.
Use enough examples to understand the task's variation. A handful of selected successes can help a team learn. They cannot establish production readiness.
6. Design accountable human review
Name the reviewer and what that person must inspect. Identify the authoritative source used to verify the output, the time available, the required expertise, the escalation route, and who owns the final action.
"Human in the loop" is not a control by itself. Review can fail when evidence is hidden, volume is too high, the reviewer lacks authority, or the interface encourages automatic approval.
For an insurer, consumer-impacting decisions require particular care. The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers describes expectations for governance, risk management, controls, and documentation under applicable insurance law. It is a model bulletin, not a universal law. Teams must check adoption and requirements in each relevant jurisdiction.
7. Keep action authority below demonstrated reliability
A pilot that produces a good draft has demonstrated drafting under tested conditions. It has not demonstrated authority to send, approve, pay, deny, bind, price, publish, or change a source record.
Separate recommendation from commitment. Keep irreversible or consequential actions behind appropriate approval while evidence is limited.
This boundary also improves diagnosis. When the system creates a proposed output without taking the final action, the team can compare its work with the reviewer decision and study disagreements.
8. Measure the full operating result
Measure output quality, unsupported claims, correction rate, reviewer agreement, review time, rework, escalation, failure recovery, user experience, latency, cost, and accessibility.
Include consequences outside the model. The workflow may require data preparation, system integration, training, logging, vendor management, incident response, or changes to an employee's role.
NIST's broader AI Risk Management Framework organizes work around governing, mapping, measuring, and managing risk. The structure is useful because measurement without ownership does not produce a controlled system.
9. Assign an owner and a stop condition
One person or accountable team should own the task definition, data boundary, evaluation, reviewer capacity, incident path, evidence record, change decisions, and retirement.
Set a duration and an exit condition. The pilot may expand, repeat with changes, pause, return to the original workflow, or end. Decide which evidence would support each outcome before the test starts.
A pilot without a stopping rule can become an unreviewed production process through habit.
10. Expand one boundary at a time
If the evidence supports continuation, change one meaningful dimension. The team might add a document type, user group, source, workflow step, or level of authority.
Retest the affected failures and confirm that reviewers, systems, documentation, and recovery can support the new scope.
The E004 strategy argument matters here. Adoption is constrained by the business, not only by model capability. The workflow must fit the organization's position, data, obligations, economics, and capacity to remain accountable.
A practical decision record
At the end of the pilot, write a short decision record in ordinary prose. State the tested task, period, participants, data boundary, model and product context, baseline, evaluation method, findings, failures, reviewer load, open risks, decision, owner, and next review date.
Link the evidence rather than summarizing away disagreement. A decision to stop is still useful. It prevents the organization from confusing novelty with value.
The next question is structural: can the task connect to authoritative data, identity, state, rules, review, action boundaries, evidence, recovery, and change ownership? [[AI Integration Requires Structural Workflows]] addresses that transition.
For the episode-level argument behind this method, read [[What Gemini 1.5 and Sora Revealed About AI Adoption]]. For a related discussion of testing products under real conditions, continue to E119.
Editorial and authority note
This guide is educational and cross-sectoral. It is not legal, regulatory, security, privacy, insurance, employment, procurement, accessibility, or investment advice. Requirements depend on the organization, tool, data, workflow, jurisdiction, and consequence. Security, privacy, legal, governance, insurance, editorial, accessibility, and founder review remain required before publication or operational use.
AI assisted with research, structure, drafting, and validation. Dalton Anderson remains the attributed author and final editorial authority.
Sources
Follow the evidence.
- Holland and Kavuri, HICSS-56aisel.aisnet.org
- Google: Our next-generation model, Gemini 1.5blog.google
- Liu et al.: Lost in the Middleaclanthology.org
- OpenAI: Video generation models as world simulatorsopenai.com
- Google AI for Developers: Long contextai.google.dev
- NAIC: Artificial Intelligencecontent.naic.org
- Spotify episode recordpodcasters.spotify.com
- OpenAI: Sora is hereopenai.com
- NAIC: Model Bulletin on the Use of Artificial Intelligence Systems by Insurerscontent.naic.org
- NIST: Artificial Intelligence Risk Management Framework, Generative Artificial Intelligence Profilenist.gov