Back to the episode map

Research Note

AI Image Editing Evaluation Research Note

AI image editing should be evaluated as two linked problems: completing the requested change and preserving content that the request did not authorize the system to chang

Aug 4, 20262 min readBy Dalton Anderson

AI Image Editing Evaluation Research Note

Editorial conclusion

AI image editing should be evaluated as two linked problems: completing the requested change and preserving content that the request did not authorize the system to change. Multi-turn evaluation adds a third problem, accumulated drift across the sequence.

No single similarity metric can replace human review of semantically important details.

Research map

SourceUseful contributionBoundary
EditInspectorHuman-annotated dimensions include edit accuracy, artifacts, visual quality, integration, common sense, and difference descriptionGoogle Research benchmark; automated evaluators can struggle and hallucinate descriptions
GIE-BenchSeparates functional correctness from preservation outside the intended edit regionResearch benchmark results do not predict every product task
CompBenchBreaks complex instructions into location, appearance, dynamics, and object requirementsBenchmark task design is broader than a single practical rubric
Google's original launch recordEstablishes first-party claims about likeness, local changes, and multi-turn editingVendor claim, not independent performance evidence
OpenAI's current image recordProvides a current comparison case for precise edits and preservation across revisionsVendor claim reviewed on July 27, 2026

Practical evaluation model

The Venture Step rubric separates instruction success, edit locality, subject identity, defining attributes, geometry and pose, text and symbols, scene preservation, integration, and artifacts.

The evaluator writes an invariant contract before seeing the first output. Each turn is compared with its immediate parent and with the original source. A critical error remains visible outside the average score.

Rights and use boundary

An image benchmark cannot establish consent, legal identity, authorship, ownership, authenticity, or permission to publish. The source, prompt, unmodified outputs, product context, rights record, and human review should remain together.

The E082 source images and outputs are missing. The transcript can motivate a method but cannot support retrospective scoring.

Sources

Follow the evidence.

  1. The effect of word concreteness on recognition memorypubmed.ncbi.nlm.nih.gov
  2. Android public naming changeblog.google
  3. GIE-Benchalphaxiv.org
  4. Systematic review of human and AI co-creativityarxiv.org
  5. ai.google.dev: image generationai.google.dev
  6. CompBenchcomp-bench.github.io
  7. Bard becomes Geminiblog.google
  8. Nano Banana across Google productsblog.google
  9. EditInspectorresearch.google
  10. Nano Banana in Google Photosblog.google
  11. Gemini 2.5 Flash Image model pageai.google.dev
  12. Nano Banana examplesblog.google
  13. Google AI updates from November 2025blog.google
  14. Canva Magic Layerscanva.com
  15. Xbox One X Project Scorpio Editionnews.xbox.com
  16. How Nano Banana got its nameblog.google
  17. Gemini app updated image editing modelblog.google
AI Image Editing Evaluation Research Note