Guide
How to Evaluate AI Image Consistency
Test AI image consistency across identity, geometry, materials, views, context, text, edit preservation, repeated runs, rights, and provenance.
In this article
How to Evaluate AI Image Consistency
AI image consistency is not one score. A generator can preserve color while changing geometry, preserve a face while changing age cues, preserve a product while inventing a logo, or make the requested edit while quietly altering everything around it.
Evaluate a defined set of protected attributes across planned views, scenes, edits, and repeated runs. Keep every output, not only the favorite, and decide whether the observed failure rate is acceptable for the intended use.
Begin with rights and intended use
Use a synthetic object you designed, a fully licensed asset, or an original reference for which the people and rights holders have explicitly authorized the input, transformation, evaluation, and intended release.
Do not use a random person's image, a private setting, a client asset, a trademarked product, a confidential design, or a dataset with uncertain rights simply because it is technically accessible.
For a person, permission should address likeness, source images, generated variations, sensitive contexts, audience, duration, withdrawal, and whether the material may be used to improve a model. For a product, record ownership of the design, marks, packaging, photography, and any third-party components.
The test does not create rights. It only evaluates behavior on an authorized input.
Write a subject contract
A subject contract describes the attributes that must remain stable and the attributes the prompt may change.
For a fictional transparent rover, protected attributes might include six wheels, equal axle spacing, a clear cabin shell, orange suspension elements, a continuous blue light strip, no logos, and a fixed wheelbase-to-height ratio. Allowed changes might include camera angle, background, lighting, and wheel rotation.
For a human subject, the contract requires more care. It might include identity, approximate age presentation, skin tone, hair, assistive device, body proportions, and wardrobe. The evaluator should also name harmful changes to watch for, including sexualization, disability removal, racialized drift, age shift, or stereotyped context.
flowchart LR
A["Rights-cleared reference"] --> B["Protected and changeable attributes"]
B --> C["Views, scenes, edits, and repetitions"]
C --> D["Preserve every output"]
D --> E["Score dimensions separately"]
E --> F["Analyze failures and intended-use risk"]
Build the test matrix before generating
| Test family | Example | What it reveals |
|---|---|---|
| View | Front, side, rear, three-quarter, overhead | Hidden geometry and contradictory structure |
| Scene | Studio, street, interior, low light | Scale, reflections, contact, shadows, and context sensitivity |
| Pose or state | Open door, moving wheel, seated person, raised hand | Articulation, body structure, and subject preservation |
| Controlled edit | Change only background color | Collateral changes to protected attributes |
| Occlusion | Partial cover, crop, object in front | Whether the model invents or substitutes hidden structure |
| Text and symbol | Fixed short label and no-logo condition | Spelling, layout, invented marks, and legibility |
| Repetition | Same prompt and settings across several runs | Output variance rather than one selected success |
| Round trip | Edit an output and then request restoration | Whether changes accumulate or protected attributes return |
Use a small matrix that you can complete. A disciplined set of 20 outputs is more useful than 200 untracked generations.
Freeze the run record
Record the provider, product surface, model identifier, model version or lifecycle label, date, region, account or project type, API or interface, prompt, system instruction, reference files, aspect ratio, resolution, seed if exposed, safety settings, tools, and every other selectable control.
Google's current Gemini API image-generation guide documents distinct image models with different reference-image limits and capabilities. It also says requested output counts may not always be followed and that generated images include SynthID. The guide was last updated July 16, 2026 at the time of this review.
Those statements apply to the documented models and interface. They do not prove that Gemini Apps, AI Studio, or a different model used the same configuration.
If the surface hides the model or seed, record that the value was unavailable. Reproducibility begins with honest missing data.
Run more than once
One output can establish that the system produced one output. It cannot establish stability.
Repeat each important condition several times while holding the visible controls constant. Keep failed generations, policy refusals, missing outputs, timeouts, and malformed files.
Do not replace a failed run with a better prompt and count the replacement as the same test. Prompt revision creates a new condition. Preserve both.
The sample size needed for a production decision depends on consequence, expected volume, cost of failure, diversity of input, and the type of harm. This guide does not declare a universal pass rate.
Score dimensions separately
| Dimension | Questions |
|---|---|
| Identity | Is the intended person or subject still the same rather than merely similar? |
| Geometry | Are proportions, parts, count, placement, and articulation coherent across views? |
| Material and color | Are finish, texture, transparency, pattern, and palette stable where protected? |
| View and pose | Does the subject remain coherent under rotation, movement, crop, and occlusion? |
| Context | Do scale, light, shadow, reflection, contact, and background relationships make sense? |
| Text and symbols | Are required words correct and prohibited marks absent? |
| Edit preservation | Did only the requested attribute change? |
| Harm and representation | Did the model introduce bias, sexualization, disability erasure, age shift, or a sensitive context? |
Use a short anchored scale. "Pass" can mean all protected attributes remained within the written tolerance. "Review" can mean the variation may be correctable. "Fail" can mean a protected or harmful attribute changed.
Write examples for each anchor before scoring. Otherwise, evaluators will redefine consistency while looking at the outputs.
Separate human judgment from ground truth
Some attributes can be measured. A wheel count, label spelling, palette value, aspect ratio, or known product dimension has a clear reference.
Others depend on human judgment. Identity, style, realism, appropriateness, and harmful representation can vary by reviewer and context. Use more than one reviewer when the decision is material. Record disagreements rather than averaging them away.
The NIST Generative AI Profile recommends defining tasks, documenting assumptions and content lineage, and evaluating accuracy, quality, reliability, and authenticity against known ground truth with multiple methods. It is a voluntary risk-management resource, not a certification of this test.
Analyze failure patterns
A production decision should be based on the distribution of failures, not the existence of a best frame.
Look for systematic drift under side views, darker skin tones, assistive devices, small text, partial occlusion, reflective materials, or repeated edits. Determine whether the failure is visible before release, repairable without changing the protected subject, and likely to recur at scale.
Manual correction can be a valid workflow if rights, time, skill, and quality control are available. It should be recorded as part of the process rather than credited to the generator.
If a model repeatedly changes a protected attribute, narrow the intended use or reject the workflow. Better prompting is not a duty to continue indefinitely.
Review provenance without overstating it
The current Google guide says generated images include SynthID. A watermark can contribute to identifying some generated media under a particular implementation. It does not prove that the depicted event happened, that the inputs were licensed, that a person consented, that the final use is fair, or that every transformation preserved the signal.
The C2PA specification site maintains the current Content Credentials standards for signed provenance assertions and asset integrity. A valid credential can help inspect declared history under a trust model. It is not proof that every assertion is complete or that the content is true.
Preserve the original inputs, prompts, model record, outputs, edits, approvals, and final exported file. Validate any credential or watermark on the final delivery surface rather than assuming it survived editing and platform transformation.
Decide by intended use
Exploration can tolerate failures that production cannot. A private mood board, internal concept, public advertisement, product representation, news image, medical illustration, and licensed character each carry different consequences.
The decision should say which uses are allowed, which need manual reconstruction, which require human subjects or rights-holder approval, and which are prohibited. It should name an owner and a refresh trigger for model, policy, or workflow changes.
A test that passes for internal ideation does not automatically pass for publication.
[[How to Review AI-Generated Assets Before Publishing]] provides the final release gate. [[What E061 Saw in Google's Shift From Chat to Canvas]] preserves the dated examples that motivated this framework.
Editorial note
This evaluation guide is not a claim that any model preserves identity, brands, products, or rights at a particular rate. No new image-generation test was run for this draft. Current model claims were checked against official Google documentation on July 28, 2026. The draft was developed with AI assistance from the preserved episode and cited sources, then prepared for image-evaluation, rights, consent, bias, accessibility, provenance, and human editorial review. Publication has not been authorized.
Sources
Follow the evidence.
- policies.google.com: use policypolicies.google.com
- open.spotify.com: 4O0DCv9Na8StBZnoXJlZ1bopen.spotify.com
- workspaceupdates.googleblog.com: introducing canvas for the gemini appworkspaceupdates.googleblog.com
- Gemini Apps Privacy Hubsupport.google.com
- ai.google.dev: image generationai.google.dev
- blog.google: gemini collaboration featuresblog.google
- daltonanderson.ghost.io: googles new ai gemini canvas consistent image modelsdaltonanderson.ghost.io
- youtu.be: qoGIyz0azwwyoutu.be
- support.google.com: 16047321support.google.com
- ai.google.dev: modelsai.google.dev
- Gemini API changelogai.google.dev
- blog.google: google gemini ai update december 2024blog.google
- w3.org: WCAG22w3.org
- NIST Generative AI Profilenvlpubs.nist.gov
- spec.c2pa.org: aboutspec.c2pa.org
- daltonanderson.net: googles new ai gemini canvas consistent image modelsdaltonanderson.net
- copyright.gov: aicopyright.gov
- cheatsheetseries.owasp.org: Secure Code Review Cheat Sheetcheatsheetseries.owasp.org