What a Published Evaluation Can Tell You
Evaluating a music generator means distinguishing two questions that are often conflated: Does the audio sound enjoyable or convincing? And does the song do what was requested? A yes to the first question does not prove a yes to the second. A polished production may still omit the requested bridge, switch instruments halfway through a section, or ignore a lyric constraint. Conversely, a song may meet several requirements without being the listeners’ preferred version.
Lyria 3.5’s official model card is the starting point for understanding which evaluations the provider reports. According to the available description, it includes human and automated evaluations, in- and out-of-distribution tests, and aspects related to prompt adherence and model limitations. This helps frame questions about what was measured, but it is not enough on its own to reconstruct an experiment: consult the current details of the procedures, evaluation instructions, sample size and composition, and generation conditions.
Google’s announcement about Lyria 3.5 also describes claimed changes in musicality, lyrics, vocals, adherence, and control over tempo and duration. These are statements by the provider about the system; they are not equivalent to an independent measurement and do not establish that any particular instruction is followed at a specific rate. To interpret them, you need to know how each capability was operationalized and whether the results apply to the channel and version you intend to use.
The ITU’s BS.1534 recommendation can inform the design of subjective audio evaluations, but its purpose is not to directly measure whether a model followed a musical prompt. Similarly, work such as MusicEval and MusicGen offers useful precedents for thinking about human ratings, quality, and alignment. It is not direct evidence of Lyria 3.5’s performance. This article’s protocol is a methodological proposal: it does not attribute results to Lyria that the available sources do not document.
Define the Test Unit Before Generating
The basic unit should not be a vague notion of “a good song” or “a prompt.” Instead, record an instruction, a generation produced under a specific condition, and an explicit list of assessable requirements. If you generate several versions from the same instruction, each audio file is a separate observation. The set can be used to study variability, but it should not become a selection of only the best examples.
For instruction-following to be verifiable, translate general expressions into observable criteria. “Make it sound cinematic” can remain a subjective rating, but it is not a binary test without an agreed definition. By contrast, “include an instrumental bridge between the second chorus and the ending” makes it possible to assess presence and placement, provided the audio has an identifiable structure. Even then, if the structure is ambiguous, evaluators should be able to mark it “undeterminable” rather than being forced to choose.
Each record should identify the exact model name and version, access channel, execution date, literal prompt, and available settings. If the channel lets you set or record parameters that affect generation, retain them. Google AI’s developer documentation identifies the stable model as `lyria-3.5`; this identifier is an example of the information to record, not a guarantee that every channel exposes identical controls.
Also decide in advance what counts as partial compliance. If an instrument is requested “throughout the song,” a brief appearance should not be recorded as full compliance. For structural requirements, distinguish the presence of a section, the order of sections, and continuity between them. Otherwise, two evaluators may give the same overall label for entirely different reasons.
Minimum test record
- 01Keep the literal prompt and break it down into independent requirements.
- 02Record the model, channel, visible version, date, and available parameters.
- 03Save every generation, including those that fail to comply or are low quality.
- 04Assign each audio file an identifier and conceal the experimental condition from evaluators.
- 05Record compliance for each requirement, confidence in the decision, and the reason for any disagreement.
Design Simple, Compound, and Conflicting Prompts
Simple tests help establish whether an isolated attribute can be evaluated consistently. For example, ask for an instrumental piece with light percussion, or a song with a lead vocal and a serene mood. These prompts do not demonstrate general control; they test one specific condition. To avoid a favorable result based on an overly broad interpretation, specify what is necessary without adding adjectives you do not intend to score.
Compound prompts test whether several requirements can coexist. One instruction might ask for a song with a verse, chorus, and bridge; a bass line present in all three sections; a moderate tempo; and lyrics that do not mention brands. Score each element independently. This makes it possible to see whether the model meets the structural requirements but fails the lyric constraint, rather than assigning a general impression such as “mostly faithful.”
Define out-of-distribution cases operationally: use less common instructions, unusual combinations, or ways of specifying an attribute that do not appear in the main test set. Do not call them “generalization” tests without explaining what makes them different. Present and analyze them separately as a stress-test set; do not mix them into the ordinary test to raise or lower a figure without context.
Conflicting instructions also deserve their own category. For example, simultaneously asking for a song “without vocals” and sung lyrics creates incompatible requirements. The result should not be scored as if there were one correct answer. The evaluation can ask whether the system identifies the conflict, prioritizes one instruction over the other, or produces an ambiguous output, but the scoring rule must be announced before listening. Without a rule set in advance, an evaluator may adapt the criterion after seeing the result.
Keep the Dimensions Separate in the Results
A useful report preserves at least four families of results. The first is instruction-following: whether each requirement appears, to what degree, and with what evidence. The second is perceived musical quality, which may include ratings of coherence, composition, or enjoyment, provided these are defined as subjective judgments rather than prompt checks. The third is audio technical quality, such as whether there are audible cuts or distortion. The fourth is vocal quality when vocals are present: intelligibility, stability, and perceived suitability are different questions from whether a voice was included when requested.
Avoid collapsing these families into a single score. An aggregate average can hide the important case where a song receives high enjoyment ratings but fails half of the structural requirements. It can also hide that two audio files with the same overall score fail on different attributes. Reporting distributions and results by requirement makes such patterns easier to identify without implying that one dimension automatically compensates for another.
Blind listening reduces some sources of bias: raters should not know which prompt produced each audio file, which generation the team prefers, or which hypothesis is being tested. Vary the listening order and keep the scoring instructions the same. If evaluators see the lyrics or prompt, the test is no longer blind with respect to that information. When that information is necessary for scoring, specify precisely what was concealed and what was not.
A rubric needs concrete anchors. For instrument presence, it might distinguish “absent,” “detectable only briefly,” “present in part of the song,” and “present in the specified sections.” For structure, it might record each section and its order, as well as an indeterminate option. These anchors are a proposal that should be tested with evaluators; they are not a universally validated scale for all music.
What to measure—and what not to infer
| Dimension | Possible observation | What is not sufficient evidence on its own |
|---|---|---|
| Instruction-following | Presence, absence, placement, or continuity of each requirement | That the song is enjoyable |
| Perceived musical quality | Human ratings of composition, coherence, or enjoyment | That every requirement was followed |
| Audio technical quality | Audible problems such as cuts or distortion | That the music is musically convincing |
| Lyrics and vocals | Lyric constraints, intelligibility, or requested vocal presence | That the rest of the structure was followed |
Human Rubrics and Inter-Rater Disagreement
Blind listening does not, by itself, guarantee a reliable evaluation. Before the main test, calibrate with practice examples and check whether the definitions are being interpreted similarly. This phase helps uncover ambiguities—for example, what counts as a “continuous instrument”—rather than eliminating legitimate differences in musical taste. Compliance criteria should refer to observable evidence, not to whether an evaluator enjoys the song.
Have more than one person rate each attribute independently, and retain the individual ratings. The report should show agreement or disagreement by dimension, not just a single overall figure. It may report agreement proportions and a concordance measure appropriate to the scale type, with an explanation of how that measure is defined. If the dataset is small or some categories occur very rarely, communicate that limitation instead of presenting estimates with an appearance of precision.
Disagreement is methodological information. If some listeners detect the bridge and others do not, perhaps the audio contains a transition that is difficult to delimit, or the rubric does not specify what counts as a bridge. If they disagree about mood, the attribute may depend on preferences or cultural references that the protocol did not define. Do not automatically resolve every difference by majority vote and then hide it: adjudication may be useful, but it should be reported separately from the original ratings.
The ITU’s BS.1534 recommendation addresses subjective audio-quality evaluation under controlled conditions. It can inspire careful preparation and presentation of stimuli, but it should not be cited as validation for a musical instruction-following rubric. “Does it sound good?” and “Is the instrument present in the specified sections?” require different questions, anchors, and analyses.
Automation: Bounded Checks, Not a Substitute for Listening
Automated analysis can be useful when the question is narrow and the method has sufficient validity for it. Duration can be checked against a target if measurable audio and a defined tolerance are available. The presence of vocals or approximate detection of certain traits can be explored with suitable tools, but any tool can produce false positives and false negatives. A detector’s label does not, on its own, prove that a person would perceive the same attribute.
Tempo, instruments, and structure are harder to turn into universal checks. An estimator may suggest beats per minute without showing that the music is perceived as having the requested character; timbre detection does not necessarily establish which instrument a listener hears; and changes in intensity do not automatically identify a verse, chorus, or bridge. Responsible technical evaluation should describe the method, its range of validity, its uncertainty, and the errors observed against human annotations.
Lyric constraints can be checked through transcription and search only if the lyrics are recognized accurately enough. A word’s absence from a transcript does not prove that it is absent from the audio. If a tool suggests a match, listen to the segment and verify its context before marking compliance or a violation. In particular, an automated analysis with no demonstrated error rate cannot prove that a ban on brand names was followed.
The point is not to exclude automation, but to use it as complementary evidence and state what it supports. MusicEval studies automated evaluation of generated music and its relationship to expert ratings; that kind of research can inform caution, but it does not validate a particular tool for measuring Lyria 3.5 in advance. If a metric has not been tested against the specific task, present it as an exploratory signal, not definitive proof.
Validity checks for an automated metric
- 01Define the exact attribute to be measured and the tool’s scope.
- 02Compare its results with independent human annotations on a relevant sample.
- 03Record false positives, false negatives, and cases where the tool cannot decide.
- 04Use automated output as support when valid, not as a replacement for listening.
- 05Describe uncertainty and avoid inferring musical quality from a technical detection.
Variability, Difficult Cases, and Reporting Results
A single generation cannot estimate model consistency. For each prompt, decide in advance how many generations to run and apply the same plan to every condition. There is no universal number to recommend without knowing the cost, expected variability, and purpose of the study. Justify and publish the decision. Selecting only the strongest outputs biases the result toward a selection capability, not the likelihood that an ordinary generation will comply.
Report compliance by requirement and prompt, as well as variation across generations. For structure, for example, show what proportion of audio files contain the requested sections and how many follow the specified order. For instruments, distinguish initial presence from continuity across the specified parts. Missing or indeterminate data should not silently be turned into successes or failures.
Compare results for compound prompts with those for simple prompts carefully. If longer instructions fail more often, this may reflect difficulty combining requirements, but it could also reflect greater textual ambiguity, interaction between attributes, or a difference in the difficulty of the test set. Interpretation should be grounded in the records and should not attribute a cause that the design cannot distinguish.
The report should include representative examples of successes, failures, and disagreements, rather than selecting only the most surprising ones. An enjoyable song that omits the bridge illustrates why quality and compliance are separate axes. An automated instrument detection that listeners cannot locate illustrates why the tool does not replace listening. Neither example, by itself, establishes the model’s overall performance.
Results worth publishing
| Result | Recommended breakdown | Limitation to make explicit |
|---|---|---|
| Compliance | By requirement, prompt, and instruction type | How partial compliance was defined |
| Consistency | Variation across generations with the same prompt | Number of generations and selection rule |
| Structure | Presence, order, and continuity of sections | Cases where structure cannot be identified |
| Human evaluation | Individual results and disagreement by attribute | Training, blinding, and adjudication |
| Automation | Method used and comparison with human annotations | Unvalidated or unmeasurable attributes |
Reproducibility and Practical Decision Criteria
A protocol is more useful when someone else can reconstruct its conditions. Publishing prompts, rubrics, evaluator instructions, inclusion and exclusion rules, the model version, and the channel used makes it possible to review what was evaluated. If access or applicable terms prevent sharing audio files, explain that limitation and provide whatever can be published. Before releasing lyrics, vocals, or other materials, also check that publication and reuse are compatible with relevant rights and conditions; do not assume that every generated asset can be redistributed without restrictions.
Exact model identification matters because an evaluation depends on the version and channel, not just the product name. If you run the protocol with the stable `lyria-3.5` model through the API, record that channel and the controls available. A result obtained through one route should not be presented as a guarantee of identical behavior through another without evidence that supports the comparison.
For production teams, a practical decision should not rest on an overall score. Start with the requirements that matter to the workflow, review which attributes were measured directly, and check how much variation there is between generations. If instrumental continuity is critical, a high subjective rating does not compensate for repeatedly failing to maintain it. If lyrics have strict constraints, they require a specific check—not a general impression that the song “followed the prompt.”
The main contribution of an instruction-following benchmark is not to proclaim a model good or bad, but to make its conditions for success and modes of failure visible. For Lyria 3.5, that means describing the instructions tested, reporting compliance and perceived quality separately, retaining the generations, and treating uncertainty as part of the result. The benchmarks index can provide context for this kind of evaluation, but a reproducible test must stand on its own protocol, data, and limitations.
Open questions
- The supplied sources do not specify in detail all prompts, sample sizes, metrics, or procedures used in Lyria 3.5’s official evaluations; check the current model card before attributing specific results.
- No universal number of generations or evaluators has been established as appropriate for every study; justify it in light of the objective, cost, and observed variability.
- The validity of automated tools for detecting tempo, instruments, structure, or lyric constraints depends on the tool and the specific task, and must be checked against human annotations.
- Access conditions, available parameters, and publication rights may vary by generation channel and must be checked for each run.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction