Back to the episode map

Research Note

AI Video Evaluation and Continuity Research Note

This note supports the model-comparison and narrative-consistency guides. The protocol is designed to be reusable across vendors and versions. It does not contain a Ventu

Aug 4, 20262 min readBy Dalton Anderson

AI Video Evaluation and Continuity Research Note

Editorial use

This note supports the model-comparison and narrative-consistency guides. The protocol is designed to be reusable across vendors and versions. It does not contain a Venture Step benchmark result because the original E088 prompts, settings, model identifiers, and raw outputs were not recovered and no new matched run was authorized or executed.

Evidence record

SourceWhat it supportsLimitation
Sora 2 System CardLaunch-era first-party capability, safety, and availability claimsSora product discontinued April 26, 2026
Veo 3.1 pageCurrent channels, creative controls, audio, and first-party preference resultsVendor-selected tests and claims
Veo 3.1 Lite model cardInputs, outputs, distribution, evaluation sample sizes, and stated limitationsLite-versus-Fast comparisons are not a universal production benchmark
NIST AI RMF Measure functionDocumented test sets, metrics, context, limitations, repeatability, and independent reviewVoluntary cross-sector framework, not a video-specific scoring system
NIST Generative AI ProfileContext-specific generative-AI risk evaluationDoes not rank video models
Face Consistency BenchmarkEvidence that face consistency can be isolated as one benchmark dimensionPreprint and narrower than narrative continuity

Comparison protocol

A fair comparison starts with a production job, not a model name. The test set should include common work, edge conditions, and at least one known failure mode. Preserve prompt text, reference inputs, seed when exposed, duration, aspect ratio, resolution, audio choice, model identifier, access channel, date, retries, moderation outcome, latency, and cost.

Forced symmetry can be misleading when models support different controls. Use a matched core test for direct comparison and a native-capability track that records what each product can do under its intended workflow.

Scoring should separate prompt adherence, visual quality, temporal coherence, identity, object and spatial continuity, audio, editability, workflow fit, latency, cost, rights, provenance, safety, and failure recovery. Record both average performance and catastrophic failures.

Continuity protocol

Character identity is one continuity dimension. A usable story also preserves wardrobe, props, spatial layout, time, lighting logic, voice, causality, and intended changes between shots.

The test uses a short story bible, five controlled shots, a continuity ledger, blinded review where practical, and an edit-burden record. A model should receive credit for a requested change and a penalty for an unrequested drift.

No score in the public guide should be presented as validated until Venture Step runs and archives the test with at least two reviewers and publishes the evidence record.

Sources

Follow the evidence.

  1. deepmind.google: veodeepmind.google
  2. deepmind.google: veo 3 1 litedeepmind.google
  3. deepmind.google: model cardsdeepmind.google
  4. openai.com: sora 2 system cardopenai.com
  5. deploymentsafety.openai.com: overview of sora 2deploymentsafety.openai.com
  6. uspto.gov: copyright and ai digital replicas report part oneuspto.gov
  7. copyright.gov: Copyright and Artificial Intelligence Part 2 Copyrightability Reportcopyright.gov
  8. openai.com: creating with sora safelyopenai.com
  9. copyright.gov: aicopyright.gov
  10. openai.com: sora 2openai.com
  11. uspto.gov: name image and likenessuspto.gov
AI Video Evaluation and Continuity Research Note