A score is not a complete evaluation
Evaluating an audio generator involves more than presenting a number. To interpret a result, readers need to know which model was run, what inputs it received, which test set was used, and what procedure calculated the metric. If any of those elements is missing, a score may still be informative within a particular experiment, but it is not enough to conclude that one system is generally better or to explain which aspect of the audio improved.
This audit is limited to the documentation supplied about Stable Audio 3 and to separating what that documentation establishes from what still needs to be checked in the complete sources. The available material identifies a paper about a model family and an official repository described as supporting inference and fine-tuning. It does not include a detailed inventory of the paper’s metrics, their values, the protocol for each test, or enough documentation to verify every experimental step. It would therefore be inappropriate to attribute results or conditions to the model that have not been confirmed.
Scope matters in the other direction, too. This article does not evaluate Stable Audio 3 in a particular production workflow or compare it with a licensed library. Its question is methodological: what needs to be known to interpret a technical evaluation, and when is it legitimate to compare its results? The page on “Stable Audio 3 Evaluation” should therefore be read as a guide to examining the evidence, not as a model ranking.
Four components that should not be confused
The model is the system that generates the audio. In the documentation available here, Stable Audio 3 is described as a family of open-weight models, with small and medium variants, based on flow matching and intended to generate instrumental music and sound effects of variable duration. This general description characterizes the system; by itself, it does not identify which variant was used in each experiment or how it was configured.
A benchmark is the set of tasks, inputs, references, and rules that determines what is being tested. It may include prompts and reference audio, but readers should not assume that all benchmarks organize these elements in the same way. A metric is a calculation that summarizes a particular property of the results under that protocol: it is neither the model nor the test set. Finally, an evaluation tool or library is software used to process files and produce one or more calculations. The existence of a library does not prove that the paper used it, or that its values match the published results.
The official Stable Audio 3 repository is described as a platform for inference and fine-tuning. The information supplied does not confirm that the repository implements a library called Stable Audio Metrics or contains the paper’s evaluation code. The description provided for Stable Audio Metrics—focused on evaluating long-form, full-bandwidth, stereo music and audio generation—helps explain the scope attributed to it, but is not enough to verify which measures it contains, what files they require, or how it was configured.
An evaluation map
Use these distinctions when reading results. They are methodological categories, not a claim that every component has been published for Stable Audio 3.
| Component | Question it answers | What to identify |
|---|---|---|
| Model | Which system produced the audio? | Variant, version or identifier, and configuration. |
| Benchmark | What task was evaluated? | Prompts, references, inclusion criteria, and test split. |
| Metric | What property does the calculation summarize? | Definition, inputs, units, aggregation, and interpretive limits. |
| Tool | How was the result calculated? | Code, version, preprocessing, and execution parameters. |
What a metric can tell you—and what it leaves out
A metric’s name is not enough to interpret its result. Readers need to check which signals it compares, how it transforms them, what it returns, and which samples it is applied to. A measure based on the relationship between generated audio and a reference poses a different question from an evaluation that summarizes the correspondence between a prompt and its audio. Even when two measures are expressed as comparable-looking numbers, they may address different objectives and rely on different assumptions.
The supplied description of Stable Audio Metrics places it in the evaluation of long-form, full-bandwidth, stereo music or generated audio. That is an indication of scope, not a verified inventory of metrics or a technical specification of their requirements. For example, it does not establish that every calculation in the library requires stereo, that all its tests measure long-duration audio, or that it was the tool used to produce the paper’s results.
Every number compresses information. An aggregate score can conceal differences across prompts, durations, content types, or generation configurations. A score also does not, by itself, establish that the result is musically coherent, that a recording is suitable for a particular use, or that the system outperforms another under different conditions. Those conclusions require tests designed to support them and additional evidence—not extrapolation from a metric’s name or value.
Conditions that make a result interpretable
A comparison depends on a precise account of what was run. First, the variant and its version need to be identified. The supplied documentation distinguishes small and medium variants, and a model card for the medium variant is available; this does not show that both variants appeared in the same evaluation, nor does it establish the settings used in a particular experiment. A family name without a sufficiently specific identifier can conceal meaningful differences between weights or versions.
Inputs matter, too. A prompt should be retained exactly as used, including changes that could alter the requested content. Interpreting the output also requires the target duration and any specified generation parameters. If several generations are produced per prompt, the report should state how many, how samples are selected, and how results are summarized. If seeds or other sampling controls are involved, they should be recorded when available. These details should not be presented as known if the sources do not document them.
The audio file is part of the protocol, not a neutral detail. Sample rate, channel count, duration, and conversions can affect the input to an analysis tool. For a reproducible comparison, it is useful to preserve the generated files before and after processing, describe any conversions, and apply the same operations across systems. Without verified specifications from the paper, it is not possible to state which settings it used; these are the fields a rigorous report should make explicit.
A minimum record for documenting a run
This process suggests how to record a future evaluation or reconstruct a published one. It does not assume that the Stable Audio 3 paper provides every item listed here.
- 01Record the variant, weight identifier, and version of the available code.
- 02Save the prompts and describe how inputs and references were selected.
- 03Record generation parameters, seeds, and the number of outputs when specified.
- 04Document duration, channels, and file characteristics, along with every transformation applied.
- 05Name the evaluation tool and version, define the metric, and specify the aggregation procedure.
- 06Publish per-sample or per-group results where possible, in addition to any overall summary.
Reproducibility: available code does not automatically make an experiment repeatable
The official documentation describes the Stable Audio 3 repository as a platform for inference and fine-tuning. That functionality may be relevant to running the model, but it does not by itself confirm that a published evaluation can be repeated. Reproducing a score also requires the specific variant, inputs and references, generation parameters, audio-processing operations, and exact implementation of the metrics.
It is useful to distinguish reproducing the system from reproducing the result. With weights and inference code, it may be possible to generate new audio; that does not necessarily recreate the files produced in the original experiment. A close replication requires the same inputs and conditions, and, when sampling introduces variation, an explanation of how that variation was controlled and summarized. Recalculating a metric also requires its code and dependencies, or a sufficiently detailed description of its implementation.
The supplied information does not establish whether the paper publishes all these items, whether its reference data can be downloaded, or which parts require additional access. The appropriate check is to review the paper and associated materials item by item. If an item is not present, report it as unverified or unavailable in the documentation consulted, rather than turning that limitation into a claim about what the complete source contains.
Contamination and data separation
Interpreting a test also depends on knowing where its prompts and references came from, how they were selected, and whether controls were used to prevent evaluation examples from coinciding with training data. A match could affect how particular results are interpreted, but raising the possibility does not demonstrate that it occurred in Stable Audio 3. Any conclusion should rest on specific documentary evidence about the data and procedure.
The supplied model card for Stable Audio 3 medium includes an excerpt related to the recording set used, but the material available here does not detail its contents or demonstrate that training data are separate from evaluation data. Separately, Stability AI states that its Stable Audio 3 music models are trained on fully licensed data. That statement concerns the provider’s claim about licensing; on its own, it does not prove that the training and test sets are independent or resolve the question of possible overlap.
An audit should look for information about sample origins, inclusion criteria, data splits, and any checks for duplicates or matches. If that information is unavailable, the rigorous wording is that separation cannot be verified from the sources consulted. It is not correct to turn a lack of public evidence into proof of contamination, but neither is it appropriate to claim that there was no overlap without documented checks.
When is a comparison justified?
Comparing results across variants, implementations, or models requires defining what stays constant. Ideally, the same prompts, test set, metric and tool version, audio processing, and equivalent rules for generating and aggregating outputs are used. If a condition changes, the report should say so and assess whether that difference could explain some of the score.
Even under a shared protocol, comparisons need context: the test set defines which cases it represents, and the metric defines which property it summarizes. A favorable result on one set does not, by itself, justify a broad claim about all music, all long-form audio, or every use case. If a variant uses different parameters, or an implementation processes audio differently, presenting a single table without explaining those differences can imply an equivalence the experiment did not establish.
The information available here does not confirm which paper results are directly comparable or whether variants were run with identical configurations. It is therefore not appropriate to infer a ranking among variants or models. The official Stable Audio 3 page can help identify the product and its variants, but it cannot replace the experimental protocol needed to compare scores.
A quick decision check before comparing two results
This table summarizes methodological criteria. A negative answer does not automatically invalidate an experiment, but it limits the conclusion the evidence can support.
| Check | If it matches | If it differs or is unknown |
|---|---|---|
| Prompt and reference set | The comparison can share the same test scope. | Do not attribute every difference to the model. |
| Variant and version | Ambiguity about the evaluated system is reduced. | Do not treat family names as sufficient identifiers. |
| Parameters and number of generations | Outputs are interpreted under more equivalent conditions. | Describe the difference and limit the inference. |
| Processing and metric | Scores are more directly comparable. | Do not assume the values measure exactly the same thing. |
Conclusion: what the evidence supports and what remains to be checked
The sources supplied make it possible to identify Stable Audio 3 as a family of models with small and medium variants, based on flow matching and described as generating instrumental music and sound effects of variable duration. They also identify an official repository intended for inference and fine-tuning, a model card for the medium variant, and official documentation about the family. These materials help locate the system and guide a review of its documentation, but they are not enough to reconstruct a complete evaluation.
In particular, the available material does not allow us to rigorously list the paper’s metrics, assign results to them, verify the generation protocol, or establish the degree of reproducibility offered by the published resources. Nor does it provide enough evidence to conclude whether training and test data overlap. These are limits of the information provided, not conclusions about the full contents of the documents.
A responsible reading of any Stable Audio 3 benchmark should keep four questions separate: which variant was evaluated, which test set was used, what property each metric calculates, and how the audio was generated and processed. Before comparing a score, verify that the conditions are equivalent and that the sources allow the calculation to be repeated or at least audited. Until those details are documented, conclusions should remain limited to what the verifiable protocol supports.
Open questions
- The supplied material does not detail which metrics the Stable Audio 3 paper lists or which results it reports.
- The supplied excerpts do not allow verification of the exact variants and versions evaluated in each experiment.
- The prompts, references, generation parameters, seeds, number of generations, and processing steps used are not specified.
- It is not confirmed that the official repository implements Stable Audio Metrics or includes the paper’s evaluation code.
- The supplied information does not establish whether test prompts or data overlap with training data.
- Without consulting the complete protocol, it is not possible to establish which paper scores are comparable to one another.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction