Research Note
AI Video Evaluation and Continuity Research Note
This note supports the model-comparison and narrative-consistency guides. The protocol is designed to be reusable across vendors and versions. It does not contain a Ventu
AI Video Evaluation and Continuity Research Note
Editorial use
This note supports the model-comparison and narrative-consistency guides. The protocol is designed to be reusable across vendors and versions. It does not contain a Venture Step benchmark result because the original E088 prompts, settings, model identifiers, and raw outputs were not recovered and no new matched run was authorized or executed.
Evidence record
| Source | What it supports | Limitation |
|---|---|---|
| Sora 2 System Card | Launch-era first-party capability, safety, and availability claims | Sora product discontinued April 26, 2026 |
| Veo 3.1 page | Current channels, creative controls, audio, and first-party preference results | Vendor-selected tests and claims |
| Veo 3.1 Lite model card | Inputs, outputs, distribution, evaluation sample sizes, and stated limitations | Lite-versus-Fast comparisons are not a universal production benchmark |
| NIST AI RMF Measure function | Documented test sets, metrics, context, limitations, repeatability, and independent review | Voluntary cross-sector framework, not a video-specific scoring system |
| NIST Generative AI Profile | Context-specific generative-AI risk evaluation | Does not rank video models |
| Face Consistency Benchmark | Evidence that face consistency can be isolated as one benchmark dimension | Preprint and narrower than narrative continuity |
Comparison protocol
A fair comparison starts with a production job, not a model name. The test set should include common work, edge conditions, and at least one known failure mode. Preserve prompt text, reference inputs, seed when exposed, duration, aspect ratio, resolution, audio choice, model identifier, access channel, date, retries, moderation outcome, latency, and cost.
Forced symmetry can be misleading when models support different controls. Use a matched core test for direct comparison and a native-capability track that records what each product can do under its intended workflow.
Scoring should separate prompt adherence, visual quality, temporal coherence, identity, object and spatial continuity, audio, editability, workflow fit, latency, cost, rights, provenance, safety, and failure recovery. Record both average performance and catastrophic failures.
Continuity protocol
Character identity is one continuity dimension. A usable story also preserves wardrobe, props, spatial layout, time, lighting logic, voice, causality, and intended changes between shots.
The test uses a short story bible, five controlled shots, a continuity ledger, blinded review where practical, and an edit-burden record. A model should receive credit for a requested change and a penalty for an unrequested drift.
No score in the public guide should be presented as validated until Venture Step runs and archives the test with at least two reviewers and publishes the evidence record.
Sources
Follow the evidence.
- deepmind.google: veodeepmind.google
- deepmind.google: veo 3 1 litedeepmind.google
- deepmind.google: model cardsdeepmind.google
- openai.com: sora 2 system cardopenai.com
- deploymentsafety.openai.com: overview of sora 2deploymentsafety.openai.com
- uspto.gov: copyright and ai digital replicas report part oneuspto.gov
- copyright.gov: Copyright and Artificial Intelligence Part 2 Copyrightability Reportcopyright.gov
- openai.com: creating with sora safelyopenai.com
- copyright.gov: aicopyright.gov
- openai.com: sora 2openai.com
- uspto.gov: name image and likenessuspto.gov