Guide
Build a High-Quality AI Tutor Source Pack
Design an AI tutor source pack with a clear learning objective, authoritative documents, versions, scope, conflicts, rights, data boundaries, gaps, and test questions.
How to Build a High-Quality Source Pack for an AI Tutor
Build an AI tutor source pack by defining the learning job first, then choosing authoritative, current, permitted sources that cover that job. Record conflicts and gaps before asking the system to generate anything.
A source-grounded assistant cannot repair a source set that is outdated, unauthorized, one-sided, or missing the answer.
flowchart TD
A["Define the learning outcome"] --> B["List questions the pack must answer"]
B --> C["Select primary and explanatory sources"]
C --> D["Record versions, authority, rights, and data class"]
D --> E["Expose conflicts and gaps"]
E --> F["Test answerable and unanswerable questions"]
F --> G["Approve, revise, or reject the pack"]
Begin with what the learner must do
"Learn the policy" is too vague.
Write a capability that can be demonstrated. The learner might need to explain a process, compare two approaches, apply a rule to a new example, diagnose a known failure, or identify when to escalate.
Then write several questions the pack should answer and at least one question it should not answer.
The unanswerable question is important. It tests whether the system exposes a gap or fills it with plausible language.
Give every source a job
Primary sources should carry rules, official specifications, policy, original research, or direct evidence. Explanatory sources can add context, examples, or a more approachable path.
Do not let an easy summary silently become the authority.
| Manifest field | What to record |
|---|---|
| Identity | Exact title, author, publisher, URL or file |
| Role | Authority, original evidence, explanation, critique, or history |
| Version | Edition, effective date, revision, and access date |
| Scope | Questions the source can and cannot answer |
| Status | Current, historical, superseded, draft, or disputed |
| Rights | Permission, license, access, and sharing boundary |
| Data class | Public, internal, confidential, personal, regulated, or prohibited |
| Conflicts | Material disagreement with another source |
| Import | Copy, sync, web extraction, transcript, or other representation |
| Omissions | Missing footnotes, comments, images, appendices, or linked material |
This manifest becomes the evidence record for the notebook.
Preserve versions instead of overwriting history
If a rule changed, keep the old and new versions as distinct records. Label their effective periods.
A question about what was allowed last year may need the historical version. A current onboarding question needs the present version.
Do not upload two editions with nearly identical names and expect the system to infer the right one. Use clear titles and ask date-specific questions.
When the source is a live web page, preserve the access date and, when permitted, a stable snapshot. A current page can change without announcing which sentence moved.
Record what the import cannot see
Google's current NotebookLM source documentation says a web import uses page text rather than nested pages or embedded media. It also notes that comments and footnotes from Google files may not be imported.
That can remove exactly the material that qualifies a claim.
Compare the imported representation with the original. Check headings, tables, footnotes, captions, scanned pages, formulas, appendices, and linked definitions.
If a source cannot be represented reliably, keep it outside the notebook and build a human review step around it.
Make conflict visible
A good pack does not force every source to agree.
Record the disputed question, each source's position, the authority and date behind it, and what evidence could resolve the difference.
Ask the assistant to describe the disagreement before asking it to recommend a conclusion. Then open the cited passages.
A model may smooth uncertainty into a single clean answer. The source manifest should make that smoothing easy to detect.
Mark the data and rights boundary
Having access to a document does not always grant permission to upload, transform, or share it.
Google's current NotebookLM privacy and terms page tells users to respect copyright and describes different terms for consumer, work, school, and cloud accounts.
The U.S. Copyright Office's AI initiative provides official policy material, but it does not decide whether a particular upload is lawful. Fair use and educational purpose are fact-specific. Institutional policy, license terms, and qualified legal review still control.
Do not upload confidential, personal, regulated, licensed, or copyrighted material merely because it would make the notebook more useful.
Remove noise without deleting dissent
Remove exact duplicates, broken files, inaccessible scans, accidental exports, and unrelated material.
Keep a source that materially disagrees with the rest when it is credible and relevant. Labeling the disagreement is better than building a falsely unanimous pack.
Separate background from decision material. A beginner explanation can help a learner form vocabulary, while the current official record remains the decision authority.
Test the pack before generating study material
Use a small acceptance set.
Ask a straightforward question with direct support. Ask a question that requires two sources. Ask a question with conflicting answers. Ask a question whose answer is absent. Ask a date-specific question where old and new versions differ.
For each response, open the citations and classify the result as supported, qualified, incomplete, contradicted, or unsupported.
The NIST AI Risk Management Framework offers a useful general pattern: map the use and affected people, measure behavior, manage the risk, and assign governance.
Use the pack as a maintained asset
Assign an owner and refresh trigger.
The trigger could be a policy revision, new course term, product update, regulation change, source correction, access change, or incident. Record which notebooks use the changed source.
Do not leave a shared AI tutor answering from a pack whose owner has moved on.
Know when the pack is insufficient
A source pack can support a bounded learning job. It should not become the final authority for medical, legal, financial, safety, employment, or other consequential decisions.
It also cannot replace a teacher, manager, records system, licensed database, or subject expert when the job depends on their authority.
Use [[How to Verify an AI Answer Against Its Citations]] to audit the generated answer. Use [[How to Evaluate an AI Tutor for a School or Workplace]] before turning a personal notebook into an institutional service. For another lesson about putting constraints in the right layer, see [[How to Write Project Rules for an AI Coding Agent]].
This guide was developed with AI assistance from the immutable E035 transcript, current NotebookLM source and privacy documentation, NIST AI RMF, U.S. Copyright Office material, and the linked source-pack framework. Dalton Anderson remains the author. Copyright, privacy, security, accessibility, education, legal, current-source, and founder review are mandatory before publication. Publication is not authorized.
Sources
Follow the evidence.
- support.google.com: 16322204support.google.com
- support.google.com: 17003757support.google.com
- support.google.com: 17004255support.google.com
- blog.google: notebooklm new features december 2024blog.google
- studentprivacy.ed.gov: privacy and education technologystudentprivacy.ed.gov
- support.google.com: 16212820support.google.com
- NIST AI Risk Management Frameworknist.gov
- www2.ed.gov: ai reportwww2.ed.gov
- daltonanderson.net: googles ai tutor the future of personalized learningdaltonanderson.net
- support.google.com: notebooklmsupport.google.com
- support.google.com: 16213268support.google.com
- journals.sagepub.com: fulljournals.sagepub.com
- youtu.be: 6BwWKkZ7aeAyoutu.be
- support.google.com: 16179559support.google.com
- NIST Generative AI Profilenvlpubs.nist.gov
- blog.google: notebooklm audio overviewsblog.google
- support.google.com: 16164461support.google.com
- open.spotify.com: 1YXy6yyC4u2mnqeARVB5qvopen.spotify.com
- edu.google.com: ai notebooklmedu.google.com
- support.google.com: 16215270support.google.com
- daltonanderson.ghost.io: googles ai tutor the future of personalized learningdaltonanderson.ghost.io