Ilustración editorial para Edición localizada de imágenes con IA: cómo comprobar que el modelo cambia solo lo pedido y conserva el resto del activo
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

A local edit is not the same as demonstrated preservation

Changing a background, removing an object, or extending a frame may appear to be bounded operations. However, generative systems produce a new output from the image, the instruction, and, where available, a mask or reference. A visually appealing result does not demonstrate that the model preserved unchanged the elements it was not asked to modify.

This distinction matters especially in e-commerce, editorial publishing, product design, and brand communications. A background edit can introduce differences in a product’s color, redraw text on packaging, alter a logo, change the shape of a hand, or modify a face. In other cases, it may shift compositional elements, vary meaningful shadows, or replace documentary details with plausible but incorrect content.

OpenAI’s image API documentation states that a mask guides the editing process, but warns that its shape may not be followed precisely. This documented capability should be understood as an intent control, not as a guarantee that every pixel outside the indicated area will remain identical. Quality control must therefore assess two separate questions: whether the authorized change occurred and whether protected areas were preserved.

Localized editing should be approached as an acceptance problem, not as a demonstration of creativity. Before generating an output, the team defines the limits of the operation and the evidence needed to approve it. If it cannot define what must be preserved, it also cannot consistently verify that preservation occurred.

This approach applies both to workflows using desktop tools and to API integrations. It is also provider-independent: it can be used when evaluating an operation with models such as GPT Image 2.5 Sunburst, Nano Banana 2, or Stable Diffusion 3.5 Large, provided the actual capabilities of the access channel in use are documented and every workflow is tested with representative assets.

02

Define the editing contract before generating

The editing contract is a short, versioned record that turns a creative request into verifiable conditions. It should accompany every test case and every published asset. Its purpose is to prevent changing criteria after seeing a favorable variant and to allow another person to reproduce the approval or rejection decision.

The first component is the permitted change. It should describe the precise operation, its location, the direction of the change, and its limits. “Improve the image” is not a verifiable instruction; “replace the gray background with a uniform white background without modifying the product or its contact shadow” makes it possible to design controls. When the expected result permits multiple visual solutions, define the acceptable range rather than requiring exact aesthetic matching.

The second component is protected regions. These may be pixels, a bounded area, detectable objects, or semantic areas: product, label, headline, face, hands, document, logo, price, or legal text. It is preferable to distinguish regions that must remain visually identical from those that may undergo small tolerable variations, such as a background texture that is retained but is not considered critical.

The third component is invariant attributes. Protecting an entire region is not always sufficient. A product may allow dust removal but not changes to its silhouette, color, relative size, number of buttons, label, or geometry. A portrait may need to retain apparent identity, expression, gaze direction, or documentary features. An editorial item may protect the legibility and literal wording of a sign while allowing changes to its surroundings.

Finally, the contract must include tolerances and rejection reasons. Tolerance is set according to risk and purpose, not according to a universal number. A decorative background may allow greater perceptual difference than a product price. It must be explicit who has authority to change the threshold: for example, the catalog owner, legal team, or editorial leadership, depending on the asset type.

Minimum fields in an editing contract

FieldQuestion it answersExample criterion
Identified source assetWhich original is being used?Internal identifier, version, and reviewed rights status
Permitted changeWhat must change?Replace only the background outside the product outline
Protected regionsWhich areas cannot be altered?Product, label, price, logo, and contact shadow
Invariant attributesWhich properties must remain?Literal text, geometry, approved color, and proportion
ToleranceWhich variation would be acceptable?No difference in text; visual review for color
RejectionWhat invalidates the output?Unreadable text, altered label, changed face, or defective edge
03

Classify the operation to assign proportionate controls

Not all edits carry the same risk. Classifying them prevents a single superficial control from being applied to heterogeneous operations. Classification should reflect both the modified area and the sensitivity of what is protected. A background replacement in a decorative image does not have the same profile as the same operation on a photograph of a medicine, an identifiable person, or a product carrying regulated information.

Background edits and object cleanup are often candidates for limited automation when the foreground is simple, clearly bounded, and contains no text or critical attributes. Even then, edges, shadows, reflections, cutouts, and possible changes to the main object must be checked. Frame-extension operations require additional caution: the added area is necessarily generated and is not evidence of a historical or documentary context absent from the original.

Edits involving products, embedded text, prices, labels, faces, and hands should be treated as high-risk categories. The issue is not only that they may contain visible failures. A variant can look correct in a quick review while changing a letter, an amount, a proportion, or an identity-related characteristic. In these cases, automation may help propose variants, but publication should depend on dedicated controls and, usually, qualified human review.

Adobe documentation for Generative Fill describes manual selection with a brush, including size and hardness adjustments, to add or replace content from an instruction. It also contemplates optional use of a reference image. Photoshop documentation describes using references on an active selection to replace an area or place an object while retaining the background. These controls are useful for preparing an operation, but the organization must independently verify whether the result fulfills its preservation contract.

Initial decision matrix by edit type

OperationDominant riskRecommended destinationMinimum control
Background without text or sensitive productCutout, edges, and colorAutomatable with samplingForeground comparison and edge review
Remove a secondary objectIncorrect reconstruction of surroundingsProposal or limited automationReview of visual continuity and neighboring elements
Extend a frameInvented contextHuman reviewLabel added area and prohibit documentary use
Product with label or packagingText, geometry, and colorProposal with mandatory approvalOCR, attribute comparison, and specialist review
Price, legal text, or documentLiteral or regulatory changeOutside the generative publication workflowDeterministic editing and content validation
Identifiable face or handsIdentity, anatomy, and consentProposal with mandatory approvalHuman review and a specific usage policy
04

Prepare a reproducible test case

A test case does not begin with the prompt. It begins with a stable, identified original. Retain the source file, its version, the date it was received or created, the intended usage context, and the applicable rights or licensing decision. Technical evaluation does not by itself resolve authorization to modify, reuse, or publish the asset; that decision must be managed in the appropriate workflow.

Version the complete instruction, not only a summarized phrase. Record the model, available version, channel or tool, exposed parameters, reference images, mask or selection, and number of requested variants. If the tool does not expose one of these elements, record its absence. Do not assume that a capability seen in another interface is available, or that the same model name implies identical outputs across products or versions.

Also prepare an observable expectation. This may consist of an expected mask, a description of the area that must change, a list of invariants, and examples of outputs that would be rejected. For a background edit, mark which part of the product will be compared. For an object modification, define which spatial relationships must remain. For a sign, retain the approved transcription against which the result will be checked.

Testing should include ordinary and adversarial cases. The latter should include thin edges, transparency, reflections, hair, small text, curved labels, glossy products, overlapping hands, and backgrounds with colors similar to the object. The goal is not to obtain a general model score, but to discover when the workflow ceases to be reliable for the specific use.

Research on MagicBrush presents a manually annotated dataset for instruction-guided image editing, with one- and multi-turn situations, with and without masks, and with quantitative, qualitative, and human evaluation. Its scope does not certify a product or establish universal organizational thresholds, but it supports the idea that evaluating instruction-guided editing is multidimensional and cannot be reduced to a single visual impression.

Test-case preparation process

  1. 01Identify and freeze the source asset; retain a copy that does not enter the generative process.
  2. 02Write the editing contract, including the permitted change, protected zones, invariants, and rejection reasons.
  3. 03Create a mask, selection, or reference when the workflow supports it; record how it was created.
  4. 04Version the prompt, tool, model, visible parameters, and number of variants.
  5. 05Define automated measurements and the type of human review before generating.
  6. 06Generate variants without overwriting the original and associate every output with the test case.
05

Measure instruction compliance and preservation separately

Evaluation must produce two independent outcomes. The first is instruction compliance: was the object removed, was the background replaced, was the requested element added in the intended region? The second is preservation: did regions and attributes outside the permission remain acceptably unchanged? An output can pass the first test and fail the second.

Pixel or region comparison can help detect unexpected changes when images are aligned and the operation is intended to leave an area unchanged. However, it should not be interpreted as a complete test of semantic identity. Small compression or rendering differences can trigger alerts, while a commercially significant alteration may affect a small area and go unnoticed in a global metric.

For that reason, combine several checks. An exclusion mask makes it possible to calculate differences outside the authorized area. OCR can compare detected text with a reference transcription, but uncertain results should be treated as a signal for review, not as automatic approval. Geometric comparison can check outlines, spatial relationships, or relative dimensions of a product. Visual-similarity or embedding methods can help prioritize anomalous variants, but they do not by themselves prove preservation of regulated, textual, or identity-related attributes.

Thresholds should be calibrated using a set of examples labeled by reviewers. At a minimum, measure false approvals, false rejections, and reviewer disagreements by category. A false approval—accepting an image that violates an invariant—usually has a higher cost for sensitive assets. The acceptable rate cannot be inferred from a tool’s documentation: it depends on risk, publication channel, and the governance of the publishing organization.

Do not use the absence of automated alerts as conclusive proof. An alert should trigger review; a favorable metric can automate a decision only when the contract, testing history, and risk level explicitly justify that delegation.

06

Apply specialized controls to text, brands, people, and products

Embedded text deserves separate treatment. Its legibility, literal wording, and, where applicable, integrity of language, currency, units, warnings, and contact details should be checked. If an image contains prices, legal information, clinical data, safety instructions, or a document, the generative workflow should not be the mechanism that determines final content. An apparently cosmetic edit can alter characters and change meaning.

Logos and brands require checking both their form and their usage context. A tool may redraw a mark with subtle variation, remove an attribution, or introduce third-party signs. Visual validation does not replace review of permissions applicable to the source asset, the reference, and the publication destination. Technical provenance does not grant usage rights either.

For identifiable faces, review should check apparent identity, features, expression, gaze, skin, hair, and context. The organization’s consent and image-rights policy must also be applied. For hands, assess anatomy, fingers, contact with objects, and continuity of accessories. These controls are qualitative and purpose-sensitive; it is not appropriate to pretend that a single metric resolves them.

For products, create a critical-attributes record by family: outline, number of components, material, approved color, label, packaging, orientation, relative size, and accessories. Compare against an authorized product reference, not the reviewer’s memory. When a reference is used to help generation, continue treating it as an input subject to rights, restrictions, and logging.

Adobe documentation on references describes uses intended to replace a selected area or place an object while preserving the background. This is a composition capability that may be useful for designing tests, but it does not prove that the produced object matches a catalog, technical specification, or brand identity.

Human review proportionate to risk

  1. 01Assign low, medium, or high risk based on the asset category and publication destination.
  2. 02Allow sampling only in categories with a stable contract, sufficient historical testing, and limited consequences.
  3. 03Require individual approval for products, faces, brands, relevant text, and editorial or documentary content.
  4. 04Where possible, separate the person who generates a variant from the person who approves publication.
  5. 05Escalate ambiguous cases to the responsible content, brand, legal, or product owner; do not turn uncertainty into tacit approval.
07

Retain variants and provenance evidence without confusing them with a quality guarantee

Generation can be non-deterministic: identical requests can yield different outputs, and an approved variant may not be reproduced exactly after a tool update. Every output should therefore retain an explicit relationship with the original, the test case, and the decision made. Do not overwrite the source asset or replace an approved variant with another that looks similar without repeating validation.

The minimum record should include the source asset identifier, reference-input identifiers, editing contract, prompt, mask or selection, declared tool and version, available parameters, generation date, produced variants, control results, reviewer identity or role, decision, rejection reason where applicable, and publication destination. Retaining rejected variants helps investigate failures, provided retention complies with internal privacy and retention rules.

C2PA defines a technical provenance specification with manifests, ingredients, and actions, including mechanisms to declare whether all actions were included. It can be used to express relationships between an asset and its derivatives or between an edit and certain recorded actions. However, its presence does not by itself demonstrate that the unedited area retains the same pixels, that text is correct, or that an output meets a publication policy. Provenance evidence and quality validation are complementary controls.

Before publication, link the specific destination to the approval. An image suitable for an internal mockup may not be suitable for a catalog, campaign, or editorial archive. The decision should state the approved channel and scope, and it should be invalidated or reviewed if the asset, prompt, mask, model, or purpose changes.

Evidence that should accompany an approved output

GroupEvidencePurpose
OriginIdentified source asset and referencesRelate the output to its inputs
InstructionContract, prompt, and versionExplain the authorized change
ExecutionTool, declared model, parameters, and maskReproduce or investigate the procedure
ValidationMetric, OCR, and review resultsJustify the quality decision
ApprovalReviewer, date, scope, and destinationDefine who authorized which use
ProvenanceManifest or compatible record where availableRetain declared relationships and actions
08

Final policy: automate, propose, or block

The final decision must be explicit and reviewable. Automate does not mean a category is never inspected; it means that, under defined conditions, it may be approved through instrumented controls and sampling. Propose means the system may produce candidates but cannot publish without an authorized person. Block means the organization has decided not to use generation for that change or asset type in the channel under consideration.

A prudent policy starts with a small scope. Select low-risk operations, create representative test cases, review every output during an initial period, and record failures. Only after observing stable results and agreeing thresholds should sampling be considered. If the model, interface, configuration, or asset type changes, reassess: observed performance in one workflow should not be automatically extrapolated to another.

This guide does not support the conclusion that any provider, model, or modality guarantees perfect preservation. The available sources describe selection, mask, reference, and provenance controls, as well as the complexity of evaluating instruction-guided editing. Operational reliability for a particular case must be demonstrated through an organization’s own tests, rejection criteria, and appropriate governance.

The desired result is not a flawless appearance in an isolated demonstration. It is a defensible decision: the team can show which change it authorized, what it protected, which controls it ran, which uncertainties it found, and why a person or system approved, sent for review, or rejected the variant.

Open questions

  • The capabilities actually available, exposed parameters, and editing behavior may vary by model, version, and access channel; they must be confirmed in the environment that will be used.
  • The supplied sources do not establish a universal pixel-difference, visual-similarity, OCR, or false-approval-rate threshold that is safe for every asset and sector.
  • Automated checks can detect anomalies, but they do not by themselves guarantee the accuracy of text, identity, product geometry, regulatory compliance, or editorial suitability.
  • Provenance information may be incomplete or may not include every action; even a complete record does not by itself prove that a visual edit preserves required attributes.
  • Authorization to edit and publish depends on rights, licenses, consent, and policies applicable to the asset, references, and distribution channel; these matters are not resolved by the technical quality of the output.
09

Keep exploring

09

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction