The useful question is not which model is best overall
If you have several local model candidates, a published score or a convincing demo is not enough to tell you which one will work in your workflow. The relevant decision is more specific: which of these models can perform the task you actually want to do to an acceptable standard, using the computer and configuration you plan to use?
An acceptance test answers that question with a small set of real or representative tasks, criteria defined before you see the responses, and comparable execution conditions. For example, a team that wants to classify requests could check whether each model assigns the right category and returns an output that its system can process. The point is not to pick the answer that sounds best, but to check observable requirements.
A small sample does not turn this method into a general ranking of models. It can reveal practical incompatibilities—such as a model failing to follow a required format or making mistakes on a common kind of input—but it does not prove that the candidate is reliable for different tasks, with other users, or on different hardware. Keep your conclusions within the limits of what you tested.
The goal also differs from comparing quantizations: for this test, it is best to hold the relevant configuration fixed rather than choose between 4, 6, or 8 bits. This is not a concurrency test, which examines simultaneous requests, or a privacy audit, which investigates network traffic. If any of those questions matters to your decision, evaluate it separately.
Start by defining the task and acceptance threshold
Before downloading candidates, describe the work you want to delegate. Avoid vague goals such as “answer well” or “be smart.” Specify what the model receives, what result you expect, and which errors would make the result unusable. The closer your description is to a real workflow task, the easier it will be to design examples that distinguish between candidates.
Also define what “acceptable” means. A draft may be useful even if it needs review; data extraction that feeds directly into a database may require every critical case to have complete fields and valid formatting. These are decisions for your team, not universal properties of a model. Write them down before comparing responses to reduce the risk of shifting the standard in favor of one candidate.
Separate criteria that describe different outcomes. Content correctness, instruction-following, and format validity are not interchangeable. A response may be correct but fail to follow the required schema; another may have perfect formatting but contain an incorrect fact. Combining these into a single unexplained score loses information that could change your decision.
If more than one person will evaluate the results, agree in advance on how you will handle disagreements. For subjective tasks, it can help to have two evaluators independently rate at least some examples and compare their reasoning. If they disagree, note which rule was ambiguous and clarify it before interpreting small differences between models. There is no need to pretend a stylistic preference has an objectively correct answer.
From a general intention to a testable criterion
Adapt the examples to your task; they are not universal requirements.
| Vague intention | Testable question | Possible acceptance criterion |
|---|---|---|
| Summarize documents | Does the model include the necessary points without adding information that is not present? | The defined points are present, and no claims unsupported by the text are added. |
| Extract information | Does the model return the requested fields in the expected format? | Required fields are correct, and the output can be parsed by the agreed process. |
| Help with internal queries | Does the model distinguish available information from information that is missing? | It answers using the provided material or acknowledges that information is missing. |
Select candidates and fix the conditions
Compare candidates that make sense for the same use. Record each model’s exact identifier and available version, along with the runtime, computer, execution environment, and relevant parameters. Without those details, an observed difference might be due to the environment rather than the model, and someone else will have difficulty reproducing the test.
Keep the conditions you can control the same: hardware, runtime and its version, instruction text, input examples, context limit, generation parameters, and how you handle the output. If a model requires a different configuration just to run, document the exception and consider whether the comparison still answers the same question. Do not change parameters halfway through the test to improve a favorite’s responses without rerunning the other candidates under equivalent conditions.
For generators whose outputs can vary between runs, also record the parameters that influence generation and repeat selected cases. Repeating a case does not eliminate all uncertainty, but it can reveal that a result depends on a lucky output. For example, Ollama’s documentation describes generation parameters that can be configured in its Modelfile format; check the current documentation for the runtime you use to confirm the available settings and what they mean.
Your memory and latency records should describe the method, not just the number. Note what operation you measured, which tool you used, how many repetitions you ran, and whether the system was doing other work. A local measurement can help you decide for that computer, but it does not automatically transfer to another device.
Minimum execution record
Complete one record per candidate and keep the same values throughout the comparison.
- 01Record the model’s exact name, identifier, and version.
- 02Record the runtime, runtime version, and computer used.
- 03Copy the prompt, parameters, context limit, and any relevant generation option.
- 04Save each input, response, and repetition with a case identifier.
- 05Record duration and memory separately, including the measurement method and conditions.
Build a small set that represents the work
Do not choose only easy examples or cases you remember because a candidate handled them well. Gather inputs typical of the intended use and add situations that often cause problems. An initial sample can be small if every case has an explicit purpose; the important thing is not to present it as a statistical representation of all the work.
Include at least three kinds of input: common cases, edge cases, and cases where the model should recognize that the available information is insufficient. If the task depends on formatting instructions, include examples that let you check whether it follows them. If it depends on user-provided content, check whether the model sticks to that content rather than filling gaps with assumptions.
Prepare an expected answer or evaluation guide for each example before running the models. You do not always need to specify one exact sentence as the correct answer: for a summary, it may be more appropriate to list the facts that must appear and claims that must not be added. For structured extraction, by contrast, you may have specific values and an expected format.
Protect the quality of the test set. Avoid including confidential information if it is not necessary, remove personal identifiers that are not part of the case, and preserve the inputs so evaluators know what information was available. If you change an example after seeing the results, mark it as a new version; do not silently combine the original test with a revised one.
Evaluate separate dimensions and record failures
A practical rubric distinguishes, at minimum, task correctness, instruction-following, format usability, and critical errors. Define what counts as a critical error for your intended use: for example, an incorrect extraction that would be processed without review may have a different cost from an awkward sentence in a draft. A generic list of risks is not a substitute for your team defining what matters.
In addition to task quality, measure latency and observed memory separately. These dimensions answer different questions: whether the result is useful, how long it takes under your chosen method, and what resources the system appeared to use during that run. Avoid combining them into one score unless you explain how you weighted them and why those weights reflect a real need. In many cases, it is more useful to present results by dimension and discuss trade-offs.
Record each failure with its case, response, unmet criterion, and assigned severity. Group failures by type, such as omission, invented information, invalid formatting, or ignored instruction, when those categories fit the task. This helps distinguish a repeated pattern from an isolated mistake. Do not hide a critical failure behind a high average on easy cases.
Performance measurements need a methodological note. The llama.cpp llama-bench tool documents repetitions and statistics such as averages and standard deviations, along with separate measures for prompt processing and generation. It also describes limits to what its measurements include. This illustrates why you should state exactly what you measured rather than treating any figure as a universal measure of speed.
Results record by dimension
This template does not prescribe weights or thresholds. Define them for your intended use and keep the observations that help interpret them.
| Dimension | What to record | Decision question |
|---|---|---|
| Correctness | Successful cases, omissions, and incorrect claims | Does the content meet the evaluation reference? |
| Instructions | Requirements followed and requirements ignored | Did the model follow the stated conditions? |
| Usable format | Validity and presence of required fields | Can the next step in the workflow use the output? |
| Critical errors | Case, error type, and expected consequence | Is there a failure that rules out the candidate? |
| Latency and memory | Measurement, tool, repetitions, and conditions | Is the observed performance adequate on this computer? |
Run the test, review the results, and make a scoped decision
Run every case with the same prompt and recorded conditions. Keep the original outputs, including defective results; editing or discarding a response before evaluation makes the test less reproducible. If outputs vary, repeat selected cases using a defined procedure and save every output, not just the one that seems most representative.
To reduce bias, evaluate responses without showing the model name when feasible. If you cannot hide it, at least apply the same rubric to every candidate and record disagreements. In human review, explicit criteria and additional evaluators where possible can help identify subjective decisions; if you report evaluator agreement, explain what was compared and how you resolved differences.
At the end, decide whether to adopt each candidate for limited use, adjust it and test again, or discard it. An adjustment changes the test conditions: save the earlier version and repeat the comparison in a comparable way. The choice may depend on a blocking requirement—for example, that every required field be valid—not on which model received the best overall score.
A responsible decision can also be “none of them meets the requirements.” If failures affect essential requirements, you do not have to choose the least bad candidate just to finish the comparison. If a candidate looks suitable, the conclusion is still provisional and limited to the examples, configuration, and computer you tested.
Decision cycle
Use the result to guide your next step, not to announce a general ranking.
- 01Check that all candidates ran the same cases under the agreed conditions.
- 02Evaluate each dimension separately and flag critical errors.
- 03Review failure patterns and variation across repetitions.
- 04Compare results with the acceptance criteria defined in advance.
- 05Adopt for a limited scope, adjust and rerun, discard, or conclude that none meets the requirements.
Limits, maintenance, and related resources
A small test is vulnerable to selection bias: the examples may favor a particular writing style, domain, or input type. Results also depend on the prompt and configuration. Publish the cases you used, the criteria you applied, and the execution conditions; do not extrapolate the result to other tasks, versions, computers, or user populations without additional evidence.
Review the test if the use changes. Adding a new data source, changing the input format, or automating a response that a person previously reviewed can change which failures matter. Keep versions of the test set and rubric so you can interpret later comparisons. An old score does not guarantee that behavior remains acceptable after you change the system.
Evaluation tools can help organize tasks and metrics, but they do not decide what success means in your case. The lm-evaluation-harness project documents tasks, metrics, custom prompts, and local evaluation. It can be a useful reference or tool when it fits your needs, but a score from a standard task does not replace your own examples or demonstrate that a model meets your requirements.
This guide focuses on accepting or rejecting candidates for a specific task. To continue exploring, see the local models guide, the model comparison tool, and the discovery area. If your next question is about quantization, concurrency, or privacy, treat it as a separate decision: changing the focus without also changing the protocol can mean your comparison no longer answers the original question.
Open questions
- Results depend on the examples, prompt, parameters, runtime, and computer; no conclusions are offered about specific models.
- The appropriate number of cases and repetitions depends on the task and observed variability; this guide sets no universal thresholds.
- Runtime options and documentation may change. Check the official documentation for the version you use.
- Local latency and memory measurements cannot be directly extrapolated to other computers or conditions.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction