Evergreen
Test AI Video Character and Story Consistency
Use a five-shot story, continuity bible, controlled changes, blinded scoring, retry and edit-burden records, and archived outputs to test AI video.
How to Test Character and Narrative Consistency in AI Video
Test AI video consistency with a short multi-shot story that defines what must stay fixed, what must change, and what each shot causes next. Preserve the prompt, references, model version, settings, retries, and every output. Score identity, wardrobe, props, space, time, voice, causality, and edit burden separately.
A stable face is not the same as a coherent story.
Write a compact story bible
Define the character with a few observable attributes, not a biography. Record clothing, carried objects, voice, location, time, lighting, and the state of important props. Assign stable names to reference assets.
Then define controlled changes. If the character removes a red coat in shot three, later shots should not be penalized for showing the intended change. If the coat changes color without instruction, that is drift.
| Continuity dimension | Fixed fact | Allowed change | Failure example |
|---|---|---|---|
| Identity | Same adult character and facial structure | Expression and viewing angle | Apparent age or identity changes |
| Wardrobe | Blue shirt and red coat | Coat removed in shot three | Shirt pattern changes |
| Object | Brass key remains with the character | Key unlocks the door | Key becomes a card |
| Space | Door is left of the window | Camera moves around the room | Door relocates |
| Time and light | Late afternoon | Light dims gradually | Night appears and reverses |
| Voice | Same speaker and accent | Emotion changes | Voice identity changes |
| Causality | Opening the door reveals the courtyard | Character may pause first | Courtyard appears before the door opens |
Use five shots that expose different failures
Shot one establishes identity and the world. Shot two adds movement and a prop interaction. Shot three introduces one deliberate state change. Shot four uses occlusion or a new camera angle. Shot five requires the story to remember the earlier change and complete a causal action.
flowchart LR
A["Establish character and world"] --> B["Move and use a prop"]
B --> C["Make one intended change"]
C --> D["Occlude, cut, or change camera"]
D --> E["Recall state and complete the consequence"]
This sequence is short enough to repeat and rich enough to expose more than portrait stability.
Test matched and native workflows
For a model comparison, use the same core story and reference material where supported. Record when one product accepts ingredients, images, scene extensions, or edit controls that another does not.
Google describes current Veo access through several products on its Veo 3.1 page. Its Flow product was introduced with ingredients, camera controls, scene building, and asset management in Google's launch record. Those workflow controls can affect continuity, so a fair report should distinguish base generation from the complete product process.
The Veo 3.1 Lite model card identifies text and image inputs, distribution channels, and vendor evaluations. It does not establish performance on this Venture Step protocol.
OpenAI described Sora 2 as improving steerability, realism, physics, and synchronized audio in its System Card. The product is now discontinued, so it belongs in historical tests rather than current tool selection.
Preserve retries and selection
Do not publish only the strongest sample. Set the allowed number of attempts before generation. Archive every output and record why one was selected.
A useful policy might allow three attempts per shot with no prompt change, followed by one documented repair. The exact rule depends on production reality. What matters is that the evaluator can distinguish first-pass reliability from curation.
Record moderation refusals and errors. They are part of workflow performance, not missing data to discard.
Score preservation and controllable change
Use a five-point scale with written anchors. A four can mean the dimension is production-usable without material repair. A three can mean small drift visible on inspection. A two can mean recognizable continuity but substantial edit work. A one can mean major inconsistency. A zero can mean failure to preserve the intended fact.
Score each shot before discussing the sequence. Then score the causal chain and overall editability.
The 2025 Face Consistency Benchmark preprint proposes standardized evaluation of character faces. It supports treating identity as measurable, but face stability alone cannot evaluate wardrobe, props, space, time, action, voice, or causal order.
NIST's AI RMF Measure function calls for test sets, metrics, tools, deployment conditions, uncertainty, and limitations to be documented. Applied here, that means the score is meaningful only beside the story bible and generation record.
Measure edit burden
Record the time and operations needed to make the sequence usable. Note regeneration, compositing, masking, color correction, audio repair, continuity edits, and manual replacement.
A lower raw score with clean controls may outperform a visually stronger model that requires repeated regeneration. Production consistency is a workflow outcome, not a single-model aesthetic trait.
Review with more than one perspective
Use at least two reviewers before treating the test as a benchmark. One should understand the intended story. Another can review the outputs with provider identity hidden where practical.
Record disagreement rather than forcing false precision. If one evaluator sees the same character and another does not, that uncertainty matters.
The protocol itself has not yet been validated by a completed Venture Step lab run. The E088 raw prompts, settings, model identifiers, and generated outputs were not recovered. This page therefore provides a reproducible test design and makes no score claim about Sora 2, Veo, or another model.
Archive the story bible, source assets, prompts, settings, outputs, refusals, score sheets, edit log, model cards, and review date. When the model changes, rerun the same sequence and preserve both records.
This page provides an evaluation protocol, not a current model ranking. It reflects sources reviewed on July 27, 2026. AI assistance was used for research organization, drafting, and validation. Publication remains unauthorized.
Sources
Follow the evidence.
- deepmind.google: veodeepmind.google
- deepmind.google: veo 3 1 litedeepmind.google
- deepmind.google: model cardsdeepmind.google
- openai.com: sora 2 system cardopenai.com
- deploymentsafety.openai.com: overview of sora 2deploymentsafety.openai.com
- uspto.gov: copyright and ai digital replicas report part oneuspto.gov
- copyright.gov: Copyright and Artificial Intelligence Part 2 Copyrightability Reportcopyright.gov
- openai.com: creating with sora safelyopenai.com
- copyright.gov: aicopyright.gov
- openai.com: sora 2openai.com
- uspto.gov: name image and likenessuspto.gov