Ilustración editorial para Stable Diffusion 3.5 Large bajo prueba: cómo medir el seguimiento de prompts
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

What question does this benchmark answer?

A prompt-following evaluation asks whether a generated image contains the elements requested by an instruction and whether they appear with the specified attributes and relationships. By itself, it does not answer whether the result is beautiful, original, or suitable for every use. Nor can it establish that a model is universally “better”: it describes behavior on a particular set of tasks, under a particular configuration, and according to a particular scoring method.

For Stable Diffusion 3.5 Large, a practical question might be: how often do generations follow verifiable requirements such as including three objects, placing one to the left of another, or displaying a specific word? The proposed comparison with Stable Diffusion 3.5 Medium should be understood as a new test. The sources available here do not provide verified quantitative results or a complete protocol that would support presenting an established ranking.

A model page, software recipe, or visual demonstration can help identify an implementation or get it running. It is not a substitute for a controlled evaluation. This article therefore describes how to produce and report evidence; it does not assign Large or Medium a prompt-following score.

02

What is verified, and what remains unknown?

Stability AI’s announcement states that Stable Diffusion 3.5 Large and inference code are available. The model card hosted under Stability AI’s Hugging Face space identifies Large as a text-to-image model and describes it as a Multimodal Diffusion Transformer. These details help identify the model and its provenance, but the available excerpts do not provide measurements of instruction following.

Amazon Bedrock documentation describes an SD 3.5 Large offering with eight billion parameters and one-megapixel output. This is a detail specific to that documentation and platform configuration; it should not automatically be generalized to every implementation or treated as evidence of quality. A vLLM recipe mentions a 2.5-billion-parameter Medium variant and an 8.1-billion-parameter Large variant within the family, but identifying those sizes does not demonstrate which variant follows prompts better.

Stability AI’s repository describes itself as a reference implementation focused on inference. The existence of code makes it possible to inspect one execution path, but does not, by itself, establish the right parameters for a comparison. Likewise, informal tests shared on Reddit are anecdotal experiences: they may suggest cases to investigate, but they cannot replace a test set with criteria fixed in advance.

The methodological conclusion is limited but important: the evidence provided does not establish which model performs better at prompt following, what difference there is between Large and Medium, or what configuration produced comparable results. The absence of such data in this set of sources also does not prove that no evaluations have been published elsewhere. It means only that results not verified here should not be attributed to the models.

Scope of the available evidence

Separating technical identification from performance evidence helps prevent descriptions or demonstrations from being presented as scores.

Available materialWhat it can supportWhat it cannot establish
Stability AI announcementAnnounced availability of Large and inference codeA prompt-following score
Model card on Hugging FaceIdentification of the model as text-to-image and a description of its architectureSuperiority to Medium or measured performance
Bedrock documentationDetails described for that platform’s offeringIndependent results or a universal configuration
Informal Reddit testAnecdotal observations that may inspire test casesA blind, reproducible comparison
03

Design prompts that make requirements testable

The evaluation set should break prompt following down into observable dimensions rather than rely on an overall impression. Useful dimensions include object presence, quantity, spatial relationships, visual attributes, and requested text. Each prompt should specify which requirement will be assessed and what counts as compliance before any images are generated. If a phrase allows multiple interpretations, rewrite it or label it as ambiguous rather than resolving the ambiguity after seeing the result.

The selection should cover both simple cases and more demanding combinations. For example, one prompt could request two cups; another, a red cup to the left of a blue book; and a third, a scene with a label on which a particular word must be legible. These are examples of protocol design, not results attributed to either model. Instructions should avoid irrelevant details that make it difficult to decide which requirement failed.

It is useful to balance the categories and record how the set is composed. If there are many object-presence prompts and few text prompts, an overall score can conceal uneven performance. Exclusions should also be defined in advance: for example, whether text will be evaluated only when clearly visible or whether typographic variation is acceptable. Prompts should not be removed because they produce images that are inconvenient for an expected conclusion.

04

Fix the execution conditions

Before generation, record the exact model and weights, along with the software version, implementation, and any pipeline modifications. Resolution, number of steps, guidance parameters, numerical precision, hardware, and acceleration options should also be published. If a platform conceals some of this information, say so and limit the conclusions accordingly.

The seed and the number of generations per prompt must be documented. A single image per instruction can allow random variation to dominate the result; multiple generations make it possible to observe that variation, although they increase the cost. The protocol should determine the number of samples before results are inspected and apply the same rule to both models. For an interpretable comparison, also state whether seeds were paired across models and what that pairing means for the implementations used.

Declaring the same resolution or number of steps is not enough if Large and Medium run through different pipelines. The cleanest comparison keeps shared parameters consistent and publishes every unavoidable difference. If memory constraints require changes to precision, batch size, or other options, those differences are part of the record and may prevent attributing the result to the model alone.

Minimum execution record

Complete and publish this record for each variant before interpreting the images.

  1. 01Identify the model, weights, version, and provenance of the implementation.
  2. 02Record the prompt, category, seed, and number of generations per prompt.
  3. 03Record resolution, steps, guidance, precision, inference options, and any acceleration.
  4. 04State the hardware, software versions, and model-specific settings.
  5. 05Save generated images and associate them with their records without revealing model identity to evaluators.
  6. 06Publish configuration differences and explain how they limit the comparison.
05

Score requirements and use blind evaluators

The main scoring unit should be the verifiable requirement. For each image, evaluators can mark whether the object is present, whether the quantity is correct, whether the requested attribute appears, whether the spatial relationship holds, and whether the text is legible and matches the request. A binary scale makes counting straightforward, though an additional “not assessable” category may be needed for corrupted images or genuinely ambiguous instructions. The rules for using it should be set before reviewing results.

Human evaluation should conceal which model generated each image and randomize the presentation order. Evaluators need shared instructions, examples showing how to apply the rubric, and a way to record uncertainty. It is advisable to use more than one evaluator and publish both agreement and disagreement; frequent disagreements may indicate that the requirement is not defined clearly enough. Disagreements should not be resolved silently, and consensus should not be presented as absolute certainty.

Automatic metrics can provide support, but they are not a universal substitute for judging every requirement. Before using them, explain what they measure, which cases they apply to, and how they have been checked against a sample reviewed by people. An overall similarity measure does not, by itself, demonstrate that there are exactly three objects, that one is to the left of another, or that a word can be read correctly. If a metric has not been validated for a dimension, the report should say so rather than treat the metric as an arbiter.

A basic rubric by dimension

Keep the scoring detailed by requirement rather than reducing all failures to a single figure at the outset.

DimensionEvaluation questionRecommended record
PresenceDoes the requested object appear?Pass, fail, or not assessable
QuantityDoes the number match the request?Observed count and required count
AttributesAre the specified colors, materials, or other properties respected?A separate result for each attribute
RelationshipIs the stated position or interaction satisfied?A result for each specific relationship
TextDoes the requested word appear and can it be read?Observed transcription and match
06

Compare Large with Medium without unwarranted attribution

The comparison should use the same prompt set and rubric. To the extent permitted by the implementations, resolution, number of generations, inference parameters, and evaluation conditions should be held constant. For each requirement, report results by model and category, along with sample counts and a description of variability, rather than only a general average.

Matching values in a parameter table does not guarantee equivalent runs if the pipeline, precision, sampling options, or software differ. Every mismatch should therefore be visible. If an important difference cannot be controlled, describe the result as a comparison between two complete configurations, not as an isolated test proving that model size or variant caused the outcome.

Do not fill unrun fields with assumed values. The protocol can be published as planned, with results marked as pending. Once data are available, each conclusion should be tied to the test set, configuration, and rubric that support it. The proposed evaluation does not predict whether Large will outperform Medium at following instructions.

Deciding whether results are comparable

Use these rules to qualify the strength of conclusions, not to hide differences.

SituationRecommended treatmentPermissible conclusion
Shared prompts, rubric, and parameters; technical differences documentedReport results by model and explain residual differencesA controlled comparison under those conditions
Relevant pipeline or parameters differBreak out the configurations and avoid attributing causality to the modelA comparison between configurations, with explicit limits
Weights, versions, or sample counts are unknownDo not present a reproducible scoreAn incomplete description that does not support a firm comparison
07

Contamination, sample selection, and limitations

A benchmark can be biased if its prompts are public, have been used repeatedly in demonstrations, or resemble examples seen during system development. Without access to information about training data, it may not be possible to confirm or rule out prior exposure. Evaluators can reduce risks by writing prompts specifically for the test, restricting access until execution, and publishing the set afterward. These should be described as partial controls, not proof that contamination is absent.

Bias also arises when only the most convincing images are selected and shown. To avoid this, retain all generations specified in advance and define exclusion rules before running the test. Example galleries can illustrate successes or failures, but they do not replace complete counts. A single seed, post hoc manual selection, or a category with very few cases can create a misleading impression.

The rubric involves human decisions: what counts as a sufficiently clear spatial relationship, when text is legible, or how much detail is enough to satisfy an attribute. This is why operational definitions, independent evaluation, and reporting disagreements matter. Results should not be combined casually with scores from other datasets, configurations, or methods: when tasks and measurement rules differ, the figures may not represent the same phenomenon.

08

Benchmark reproducibility checklist

A useful evaluation should let someone else repeat the procedure and understand its limitations, even if they do not obtain identical images. Published materials should include prompts, inclusion and exclusion criteria, seeds, number of generations, versions, and configurations, as well as the images or a clear explanation of why they are not distributed. They should also include the rubric, evaluation forms, blinding method, and results broken down by dimension.

The report should distinguish observed facts, methodological decisions, and interpretations. For example, an evaluator response count is an execution result; the definition of “spatial relationship satisfied” is a rubric decision; claiming that a difference is caused by the model is an interpretation requiring comparable conditions and sufficient evidence. This separation helps readers avoid confusing a protocol recommendation with a result that has already been measured.

Until an evaluation of this kind is published and reproduced, the responsible approach is to present Stable Diffusion 3.5 Large and Medium as variants that require testing under explicit conditions, not as winners or losers in a prompt-following benchmark. The benchmarks index can provide editorial context, and the SD 3.5 Large evaluation page can serve as the reference for the test, provided pending results are not described as measurements. The practical conclusion is straightforward: publish the method before interpreting the images, and publish the data alongside any score.

Reproducible publication checklist

Check every item before presenting a score as a benchmark result.

  1. 01Complete prompts, categories, and rationale for exclusions.
  2. 02Model identifiers, weights, software, implementations, and parameters.
  3. 03Seeds, number of generations, resolution, and execution conditions.
  4. 04Generated samples with identifiers that let each evaluation be traced.
  5. 05Rubric, evaluator instructions, blinding, and rules for handling disagreements.
  6. 06Results by category and requirement, with limitations and configuration differences.
  7. 07Sufficient instructions to repeat the process and report any deviations.

Open questions

  • The sources provided do not verify a primary evaluation reporting quantitative prompt-following results for SD 3.5 Large.
  • No Large-versus-Medium comparison protocol is available that would allow differences to be attributed exclusively to the model variant.
  • Information about model sizes comes from pages and recipes with different scopes and does not establish performance results.
  • The prompt set, conditions, number of evaluators, and rubric have not been fixed or run; this article presents a proposed method, not results.
  • Possible prior exposure of the models to test prompts or images cannot be determined from the evidence provided.
09

Keep exploring

09

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction