Research Note
AI Image Editing Evaluation Research Note
AI image editing should be evaluated as two linked problems: completing the requested change and preserving content that the request did not authorize the system to chang
AI Image Editing Evaluation Research Note
Editorial conclusion
AI image editing should be evaluated as two linked problems: completing the requested change and preserving content that the request did not authorize the system to change. Multi-turn evaluation adds a third problem, accumulated drift across the sequence.
No single similarity metric can replace human review of semantically important details.
Research map
| Source | Useful contribution | Boundary |
|---|---|---|
| EditInspector | Human-annotated dimensions include edit accuracy, artifacts, visual quality, integration, common sense, and difference description | Google Research benchmark; automated evaluators can struggle and hallucinate descriptions |
| GIE-Bench | Separates functional correctness from preservation outside the intended edit region | Research benchmark results do not predict every product task |
| CompBench | Breaks complex instructions into location, appearance, dynamics, and object requirements | Benchmark task design is broader than a single practical rubric |
| Google's original launch record | Establishes first-party claims about likeness, local changes, and multi-turn editing | Vendor claim, not independent performance evidence |
| OpenAI's current image record | Provides a current comparison case for precise edits and preservation across revisions | Vendor claim reviewed on July 27, 2026 |
Practical evaluation model
The Venture Step rubric separates instruction success, edit locality, subject identity, defining attributes, geometry and pose, text and symbols, scene preservation, integration, and artifacts.
The evaluator writes an invariant contract before seeing the first output. Each turn is compared with its immediate parent and with the original source. A critical error remains visible outside the average score.
Rights and use boundary
An image benchmark cannot establish consent, legal identity, authorship, ownership, authenticity, or permission to publish. The source, prompt, unmodified outputs, product context, rights record, and human review should remain together.
The E082 source images and outputs are missing. The transcript can motivate a method but cannot support retrospective scoring.
Sources
Follow the evidence.
- The effect of word concreteness on recognition memorypubmed.ncbi.nlm.nih.gov
- Android public naming changeblog.google
- GIE-Benchalphaxiv.org
- Systematic review of human and AI co-creativityarxiv.org
- ai.google.dev: image generationai.google.dev
- CompBenchcomp-bench.github.io
- Bard becomes Geminiblog.google
- Nano Banana across Google productsblog.google
- EditInspectorresearch.google
- Nano Banana in Google Photosblog.google
- Gemini 2.5 Flash Image model pageai.google.dev
- Nano Banana examplesblog.google
- Google AI updates from November 2025blog.google
- Canva Magic Layerscanva.com
- Xbox One X Project Scorpio Editionnews.xbox.com
- How Nano Banana got its nameblog.google
- Gemini app updated image editing modelblog.google