COMPARABILITY

Evaluation methodology

We separate third-party benchmarks from tests run by Inferama and record the conditions needed to interpret results.

Last reviewed: Operated by: Linkses OÜ
01

Two kinds of result

Provider results keep their source, metric and published conditions. Inferama runs record the exact model, provider, date, prompt, parameters, tools, attempts and response.

02

Scoring and review

Each domain uses an explicit rubric. Automated judges are identified, and ambiguous or sensitive results require additional review.

03

Reproducibility and limits

A result is a dated sample, not a universal quality guarantee. Different tasks, tools or inference budgets are not directly comparable.

04

Automated execution and artefacts

Compatible text tests run in bounded provider batches with encrypted credentials, retry and daily limits. Image, video, audio and other artefact tests remain manual until a verifiable file is available and inspected.