Research Note

Bounded AI Assistant Evaluation Framework

Evaluation begins with a release decision. State the exact task, audience, data class, consequence, version, and authority being considered. A pass supports only that bou

Aug 4, 20262 min readBy Dalton Anderson
In this article

Bounded AI Assistant Evaluation Framework

Decision before score

Evaluation begins with a release decision. State the exact task, audience, data class, consequence, version, and authority being considered. A pass supports only that boundary.

Test classes

Representative cases cover ordinary work. Edge cases cover incomplete, noisy, contradictory, and unusually long inputs. Adversarial cases challenge instructions, source rules, and permission boundaries. Privacy cases introduce data that should be rejected or handled differently. Recovery cases test whether a failure can be detected, corrected, escalated, and reversed.

Observable properties

Measure task completion, source fidelity, factual support, instruction adherence, calibration, privacy, permission, consistency, recovery, and reviewer effort. Exact-match scoring is useful only when one answer is genuinely expected.

Severity

Classify failures by consequence. A harmless formatting miss is not equivalent to exposing private data, fabricating a source, or taking an unauthorized action. One severe failure can justify a hold even when the average score looks high.

Run record

Freeze instructions, sources, product surface, model if visible, tools, permissions, test data, date, and evaluator. Preserve the prompts, outputs, scores, reviewer notes, and any incidents.

Change rule

Retest when instructions, sources, model, product, tools, permissions, audience, task, or policy changes. Monitor production for failure patterns that the original set missed.

Governance basis

NIST's AI Risk Management Framework and Generative AI Profile support use-case-specific governance, measurement, pre-deployment testing, ongoing evaluation, and incident handling. They do not certify a particular assistant or replace domain review.

Sources

Follow the evidence.

  1. about.fb.com: create your own custom ai with ai studioabout.fb.com
  2. ai.meta.com: ai studioai.meta.com
  3. blog.google: google gemini update august 2024blog.google
  4. daltonanderson.ghost.io: google gems vs meta ai building your first ai agentdaltonanderson.ghost.io
  5. open.spotify.com: 0ZMJAP0X2CzPVbC83gaWagopen.spotify.com
  6. privacycenter.instagram.com: policyprivacycenter.instagram.com
  7. Gemini Apps Privacy Hubsupport.google.com
  8. support.google.com: 15235603support.google.com
  9. support.google.com: 15146780support.google.com
  10. support.google.com: 16504957support.google.com
  11. tsapps.nist.gov: get pdftsapps.nist.gov
  12. facebook.com: 1675196359893731facebook.com
  13. NIST AI Risk Management Frameworknist.gov
  14. youtu.be: nAW62 6pXaUyoutu.be

From this episode

Two useful next steps.

Guide · 1 min

How to Write Custom AI Assistant Instructions

Write testable custom AI instructions that define the job, evidence, output, limits, uncertainty, permissions, escalation, examples, and owner.

Guide · 1 min

How to Test a Custom AI Assistant Before Use

Test a bounded AI assistant with representative, edge, adversarial, privacy, permission, consistency, recovery, and human-review cases.

Return to the episode