Two kinds of result
Provider results keep their source, metric and published conditions. Inferama runs record the exact model, provider, date, prompt, parameters, tools, attempts and response.
Scoring and review
Each domain uses an explicit rubric. Automated judges are identified, and ambiguous or sensitive results require additional review.
Reproducibility and limits
A result is a dated sample, not a universal quality guarantee. Different tasks, tools or inference budgets are not directly comparable.
Automated execution and artefacts
Compatible text tests run in bounded provider batches with encrypted credentials, retry and daily limits. Image, video, audio and other artefact tests remain manual until a verifiable file is available and inspected.