Back to the episode map

Evergreen

How to Test Enterprise AI Search and Summaries

Test enterprise AI search for retrieval, claim support, citation fidelity, conflicting and stale sources, permissions, abstention, correction time, and harmful failures.

Aug 4, 20267 min readBy Dalton Anderson

How to Test AI Search and Summaries Against Sources

Enterprise AI search should be tested as an evidence system, not a writing contest. A useful answer retrieves the right material, distinguishes authority and freshness, supports consequential claims with inspectable sources, respects access boundaries, and admits when the evidence is missing or conflicting.

The test needs a sealed answer set and permission cases. Otherwise a fluent answer can pass because the evaluator does not know what the system should have found.

flowchart LR
    A["Question with known evidence state"] --> B["AI retrieval and response"]
    B --> C["Inspect retrieved sources"]
    C --> D["Map each consequential claim to evidence"]
    D --> E["Check authority, date, conflict, and permissions"]
    E --> F{"Supported and properly bounded?"}
    F --> G["Accept"]
    F --> H["Correct, abstain, or fail"]
    H --> I["Record cause and consequence"]

Define the source universe

Start by identifying the repositories the product is allowed to search. Record the tenant, plan, feature, connected systems, indexed locations, excluded locations, user role, administrative settings, and test date.

The word “enterprise” does not imply that every internal source is available. Microsoft says Microsoft 365 Copilot can use Microsoft Graph and connected data in several experiences, but its product overview also describes narrower surfaces. A Teams chat answer can be confined to one chat thread.

Slack’s AI feature guide says search answers draw on information available to the user and include citations. Enterprise search can add connected sources on the documented plan.

Test the exact product surface. Do not assume the search bar, meeting assistant, app sidebar, and general chat share the same source corpus.

Create a source authority map

Before writing questions, identify which records are authoritative. A current approved policy should outrank an old draft. A signed decision record may outrank a discussion thread. A system-of-record value may outrank a copied spreadsheet.

Record the owner, status, effective date, superseded version, and permitted audience for each controlling source.

The source map exposes an organizational problem that AI cannot solve by itself. If the company has several files claiming to be the current policy, retrieval can surface the conflict but cannot invent legitimate authority.

Keep the ambiguity in the test. A good answer should identify the conflict and avoid presenting one version as settled without evidence.

Build a sealed answer set

Write test questions before exposing the expected answers to the pilot users or system. For each question, preserve the answer state, controlling evidence, acceptable phrasing, required caveats, prohibited sources, access rule, and consequence of error.

The set should include more than obvious fact retrieval.

CaseEvidence stateExpected behavior
Direct factOne current authoritative sourceAnswer and cite the supporting passage
SynthesisSeveral compatible sourcesCombine without changing scope
Missing factNo approved evidenceSay the answer is unavailable
ConflictCurrent sources disagreeDescribe the conflict and avoid false resolution
Stale recordOld answer is easier to findPrefer current authority and identify supersession
Restricted factEvidence exists outside user accessWithhold it without revealing the protected content
Ambiguous questionSeveral interpretations fitAsk for clarification or state the chosen interpretation
Manipulated sourceA source contains instructions aimed at the modelTreat the source as evidence, not authority over the system

Use language and misspellings that real employees use. Include abbreviations, renamed projects, uncommon names, and questions that require a date or unit.

Separate retrieval from generation

An answer can fail because the system did not retrieve the right source, because it misread the source, or because it invented a connection after retrieval. Score those failures separately.

First inspect which sources were retrieved. Then map every consequential answer claim to a source passage. Finally determine whether the conclusion stays within what those passages establish.

This distinction makes remediation possible. A retrieval failure may require indexing, naming, metadata, permissions, or authority cleanup. A generation failure may require a different workflow, prompt, model, response constraint, or human review.

Open every citation

A citation is not a decorative trust signal. It is a route to evidence.

Confirm that the link opens for the user, reaches the cited version, and supports the nearby statement. Check the date, owner, status, passage, and context. A source may contain the same keywords while contradicting the answer.

Microsoft’s current overview says Teams chat responses can provide clickable citations to the source content used. Slack says its search answers include citations to source messages or files. Those are product capabilities, not proof that any particular citation supports the generated claim.

Record citation precision. A link to a long document may be technically valid while leaving the reviewer unable to locate the evidence. Note whether the system points to the passage, item, message, or only the container.

Test missing and conflicting evidence

Many demonstrations reward the system for answering. Real knowledge work sometimes requires restraint.

Ask questions whose answers are not present. Ask about a policy that has two active-looking versions. Ask for a number that appears with different units. Ask a question whose premise is false.

The correct behavior may be to abstain, name the missing evidence, describe the conflict, or ask for clarification. Penalize an answer that resolves uncertainty through confident invention.

NIST’s Generative AI Profile identifies confabulation and information-integrity risks among the considerations for generative AI. The test should make those risks observable in the organization’s own material.

Test permissions with paired users

Create paired test identities whose access differs in one intentional way. Ask the same restricted and unrestricted questions through both accounts.

The authorized user should receive the permitted evidence. The unauthorized user should not receive the protected answer, source title, revealing summary, or enough fragments to reconstruct it.

Also test oversharing. If both users can retrieve sensitive material because the underlying repository is too broad, the AI may be honoring enforced permissions while the organization is still failing its intended access policy.

Slack’s AI security record and Microsoft’s Copilot security documentation describe their respective current protections. The paired-user test verifies the organization’s actual configuration.

Score the completed answer

Report answer correctness, required-element completeness, retrieval recall for known sources, unsupported-claim rate, citation support, conflict handling, abstention quality, access behavior, correction time, and harmful failure.

Do not collapse everything into one average. A system can answer common questions well and still be unacceptable because it occasionally exposes restricted material or invents a policy requirement.

Preserve examples of the most consequential failures. The original question, user role, retrieved sources, answer, reviewer judgment, correction, and downstream risk are more useful than a vague “accuracy issue” count.

Use failures to improve the source system

Search testing often reveals that the knowledge base needs work. Important files may have unclear titles, missing dates, duplicated authority, weak access controls, obsolete copies, or no owner.

Fix those issues in the source system and rerun the same cases. Do not tune the test until the system passes.

A search assistant is downstream of organizational memory. Its quality depends on the records it can reach and the authority model those records express.

Make the decision narrow and dated

State which questions, sources, users, and consequences the test covered. Record the product configuration and date. Approve only the use that the evidence supports.

Set a retest trigger for changes to the model, plan, connected sources, indexing, permission structure, authority map, citation behavior, or workflow. [[How to Run a Bounded Workplace AI Pilot]] provides the broader operating container.

The goal is not to prove that AI can answer questions. It is to know when an answer deserves to influence work.

Editorial note

This guide was developed with AI assistance from the immutable E042 transcript and the linked Microsoft, Slack, NIST, enterprise-search, security, privacy, and records sources. Dalton Anderson remains the author. Security, privacy, records, accessibility, legal, domain, source, and founder review are mandatory before publication. Publication is not authorized.

Sources

Follow the evidence.

  1. slack.com: 28244420881555 Manage access to AI features in Slackslack.com
  2. learn.microsoft.com: recording transcription overviewlearn.microsoft.com
  3. open.spotify.com: 0FyyANPnMYdcc04GiM2OWXopen.spotify.com
  4. daltonanderson.ghost.io: ai in the workplace is copilot and slack ai worth itdaltonanderson.ghost.io
  5. learn.microsoft.com: security microsoft 365 copilotlearn.microsoft.com
  6. iea.org: key questions on energy and aiiea.org
  7. iea.org: data centre electricity use surged in 2025 even with tightening bottlenecks driving a scramble for solutionsiea.org
  8. slack.com: 31377193680019 Use AI to take huddle notes in Slackslack.com
  9. NIST AI Risk Management Frameworknist.gov
  10. iea.org: executive summaryiea.org
  11. slack.com: 115004846068 Slack updates and changesslack.com
  12. learn.microsoft.com: microsoft 365 copilot overviewlearn.microsoft.com
  13. slack.com: 28310650165907 Security for AI features in Slackslack.com
  14. youtu.be: ZMvMBflUd4youtu.be
  15. NIST Generative AI Profilenvlpubs.nist.gov
  16. slack.com: 25076892548883 Guide to AI features in Slackslack.com
  17. daltonanderson.net: ai in the workplace is copilot and slack ai worth itdaltonanderson.net