Exact version
Check the model identifier, date and whether the alias can change.
Published results and conditions declared by their sources. Figures are comparable only when version, configuration, metric and date match.
| Model | Organization | Result | Metric | Conditions |
|---|---|---|---|---|
| Stable Diffusion 3.5 Large | Stability AI | Main base model | Quality and prompt adherence | Result or evaluation published by the provider; review methodology and conditions in the source. |
Check the model identifier, date and whether the alias can change.
Review tools, effort, number of attempts, prompt and inference budget.
Validate the result with tasks, languages, formats and errors representative of your product.
A public benchmark provides guidance, but does not replace an evaluation with your data, tools, budget, and error tolerance.
A moving alias and a snapshot may produce different results. Record the model, date, and provider.
Prompt, reasoning, tools, number of attempts, and budget must be equivalent for comparison.
A small difference may disappear between runs. Keep samples, dispersion, and failures, not just the average.
Validate whether the improvement holds across languages, formats, and real cases close to your product.
A metric is only useful when we know the task, the conditions, and the error that matters. The result begins with the design of the test.