Guide

How to Evaluate AI Image Consistency

Test AI image consistency across identity, geometry, materials, views, context, text, edit preservation, repeated runs, rights, and provenance.

Aug 4, 20267 min readBy Dalton Anderson
In this article

How to Evaluate AI Image Consistency

AI image consistency is not one score. A generator can preserve color while changing geometry, preserve a face while changing age cues, preserve a product while inventing a logo, or make the requested edit while quietly altering everything around it.

Evaluate a defined set of protected attributes across planned views, scenes, edits, and repeated runs. Keep every output, not only the favorite, and decide whether the observed failure rate is acceptable for the intended use.

Begin with rights and intended use

Use a synthetic object you designed, a fully licensed asset, or an original reference for which the people and rights holders have explicitly authorized the input, transformation, evaluation, and intended release.

Do not use a random person's image, a private setting, a client asset, a trademarked product, a confidential design, or a dataset with uncertain rights simply because it is technically accessible.

For a person, permission should address likeness, source images, generated variations, sensitive contexts, audience, duration, withdrawal, and whether the material may be used to improve a model. For a product, record ownership of the design, marks, packaging, photography, and any third-party components.

The test does not create rights. It only evaluates behavior on an authorized input.

Write a subject contract

A subject contract describes the attributes that must remain stable and the attributes the prompt may change.

For a fictional transparent rover, protected attributes might include six wheels, equal axle spacing, a clear cabin shell, orange suspension elements, a continuous blue light strip, no logos, and a fixed wheelbase-to-height ratio. Allowed changes might include camera angle, background, lighting, and wheel rotation.

For a human subject, the contract requires more care. It might include identity, approximate age presentation, skin tone, hair, assistive device, body proportions, and wardrobe. The evaluator should also name harmful changes to watch for, including sexualization, disability removal, racialized drift, age shift, or stereotyped context.

flowchart LR
    A["Rights-cleared reference"] --> B["Protected and changeable attributes"]
    B --> C["Views, scenes, edits, and repetitions"]
    C --> D["Preserve every output"]
    D --> E["Score dimensions separately"]
    E --> F["Analyze failures and intended-use risk"]

Build the test matrix before generating

Test familyExampleWhat it reveals
ViewFront, side, rear, three-quarter, overheadHidden geometry and contradictory structure
SceneStudio, street, interior, low lightScale, reflections, contact, shadows, and context sensitivity
Pose or stateOpen door, moving wheel, seated person, raised handArticulation, body structure, and subject preservation
Controlled editChange only background colorCollateral changes to protected attributes
OcclusionPartial cover, crop, object in frontWhether the model invents or substitutes hidden structure
Text and symbolFixed short label and no-logo conditionSpelling, layout, invented marks, and legibility
RepetitionSame prompt and settings across several runsOutput variance rather than one selected success
Round tripEdit an output and then request restorationWhether changes accumulate or protected attributes return

Use a small matrix that you can complete. A disciplined set of 20 outputs is more useful than 200 untracked generations.

Freeze the run record

Record the provider, product surface, model identifier, model version or lifecycle label, date, region, account or project type, API or interface, prompt, system instruction, reference files, aspect ratio, resolution, seed if exposed, safety settings, tools, and every other selectable control.

Google's current Gemini API image-generation guide documents distinct image models with different reference-image limits and capabilities. It also says requested output counts may not always be followed and that generated images include SynthID. The guide was last updated July 16, 2026 at the time of this review.

Those statements apply to the documented models and interface. They do not prove that Gemini Apps, AI Studio, or a different model used the same configuration.

If the surface hides the model or seed, record that the value was unavailable. Reproducibility begins with honest missing data.

Run more than once

One output can establish that the system produced one output. It cannot establish stability.

Repeat each important condition several times while holding the visible controls constant. Keep failed generations, policy refusals, missing outputs, timeouts, and malformed files.

Do not replace a failed run with a better prompt and count the replacement as the same test. Prompt revision creates a new condition. Preserve both.

The sample size needed for a production decision depends on consequence, expected volume, cost of failure, diversity of input, and the type of harm. This guide does not declare a universal pass rate.

Score dimensions separately

DimensionQuestions
IdentityIs the intended person or subject still the same rather than merely similar?
GeometryAre proportions, parts, count, placement, and articulation coherent across views?
Material and colorAre finish, texture, transparency, pattern, and palette stable where protected?
View and poseDoes the subject remain coherent under rotation, movement, crop, and occlusion?
ContextDo scale, light, shadow, reflection, contact, and background relationships make sense?
Text and symbolsAre required words correct and prohibited marks absent?
Edit preservationDid only the requested attribute change?
Harm and representationDid the model introduce bias, sexualization, disability erasure, age shift, or a sensitive context?

Use a short anchored scale. "Pass" can mean all protected attributes remained within the written tolerance. "Review" can mean the variation may be correctable. "Fail" can mean a protected or harmful attribute changed.

Write examples for each anchor before scoring. Otherwise, evaluators will redefine consistency while looking at the outputs.

Separate human judgment from ground truth

Some attributes can be measured. A wheel count, label spelling, palette value, aspect ratio, or known product dimension has a clear reference.

Others depend on human judgment. Identity, style, realism, appropriateness, and harmful representation can vary by reviewer and context. Use more than one reviewer when the decision is material. Record disagreements rather than averaging them away.

The NIST Generative AI Profile recommends defining tasks, documenting assumptions and content lineage, and evaluating accuracy, quality, reliability, and authenticity against known ground truth with multiple methods. It is a voluntary risk-management resource, not a certification of this test.

Analyze failure patterns

A production decision should be based on the distribution of failures, not the existence of a best frame.

Look for systematic drift under side views, darker skin tones, assistive devices, small text, partial occlusion, reflective materials, or repeated edits. Determine whether the failure is visible before release, repairable without changing the protected subject, and likely to recur at scale.

Manual correction can be a valid workflow if rights, time, skill, and quality control are available. It should be recorded as part of the process rather than credited to the generator.

If a model repeatedly changes a protected attribute, narrow the intended use or reject the workflow. Better prompting is not a duty to continue indefinitely.

Review provenance without overstating it

The current Google guide says generated images include SynthID. A watermark can contribute to identifying some generated media under a particular implementation. It does not prove that the depicted event happened, that the inputs were licensed, that a person consented, that the final use is fair, or that every transformation preserved the signal.

The C2PA specification site maintains the current Content Credentials standards for signed provenance assertions and asset integrity. A valid credential can help inspect declared history under a trust model. It is not proof that every assertion is complete or that the content is true.

Preserve the original inputs, prompts, model record, outputs, edits, approvals, and final exported file. Validate any credential or watermark on the final delivery surface rather than assuming it survived editing and platform transformation.

Decide by intended use

Exploration can tolerate failures that production cannot. A private mood board, internal concept, public advertisement, product representation, news image, medical illustration, and licensed character each carry different consequences.

The decision should say which uses are allowed, which need manual reconstruction, which require human subjects or rights-holder approval, and which are prohibited. It should name an owner and a refresh trigger for model, policy, or workflow changes.

A test that passes for internal ideation does not automatically pass for publication.

[[How to Review AI-Generated Assets Before Publishing]] provides the final release gate. [[What E061 Saw in Google's Shift From Chat to Canvas]] preserves the dated examples that motivated this framework.

Editorial note

This evaluation guide is not a claim that any model preserves identity, brands, products, or rights at a particular rate. No new image-generation test was run for this draft. Current model claims were checked against official Google documentation on July 28, 2026. The draft was developed with AI assistance from the preserved episode and cited sources, then prepared for image-evaluation, rights, consent, bias, accessibility, provenance, and human editorial review. Publication has not been authorized.

Sources

Follow the evidence.

  1. policies.google.com: use policypolicies.google.com
  2. open.spotify.com: 4O0DCv9Na8StBZnoXJlZ1bopen.spotify.com
  3. workspaceupdates.googleblog.com: introducing canvas for the gemini appworkspaceupdates.googleblog.com
  4. Gemini Apps Privacy Hubsupport.google.com
  5. ai.google.dev: image generationai.google.dev
  6. blog.google: gemini collaboration featuresblog.google
  7. daltonanderson.ghost.io: googles new ai gemini canvas consistent image modelsdaltonanderson.ghost.io
  8. youtu.be: qoGIyz0azwwyoutu.be
  9. support.google.com: 16047321support.google.com
  10. ai.google.dev: modelsai.google.dev
  11. Gemini API changelogai.google.dev
  12. blog.google: google gemini ai update december 2024blog.google
  13. w3.org: WCAG22w3.org
  14. NIST Generative AI Profilenvlpubs.nist.gov
  15. spec.c2pa.org: aboutspec.c2pa.org
  16. daltonanderson.net: googles new ai gemini canvas consistent image modelsdaltonanderson.net
  17. copyright.gov: aicopyright.gov
  18. cheatsheetseries.owasp.org: Secure Code Review Cheat Sheetcheatsheetseries.owasp.org

From this episode

Two useful next steps.

Evergreen · 1 min

What Gemini Canvas Is and When to Use It

Gemini Canvas is an editable workspace inside Gemini Apps for documents, apps, slides, and code. Learn how it differs from chat, APIs, and production tools.

Guide · 1 min

How to Review AI-Generated Assets Before Publishing

Review AI-generated text, images, and code for truth, sources, rights, consent, privacy, security, accessibility, provenance, approval, and final-channel behavior.

Return to the episode