Recommendations Are a Starting Point, Not a Guarantee
A prompt template can improve a model’s answer while making integration worse. For example, an instruction that encourages more explanation may help with an analysis task but violate the contract of an API that expects a brief JSON response. When deploying DeepSeek R1, the practical question is therefore not only what configuration the provider recommends, but whether that configuration improves real tasks without introducing product-affecting regressions.
DeepSeek’s official R1 repository provides usage recommendations concerning where instructions are placed, the system message, and how the assistant’s response begins. It also recommends running multiple tests during evaluation. These guidelines offer a basis for designing experiments, but they do not by themselves prove that any one practice is superior in every application. The team defining the tasks, acceptance criteria, and consequences of each error is responsible for validating them.
This article proposes a controlled ablation: change one factor at a time, keep the others constant, and record not only whether an answer seems correct but also whether it satisfies the output contract. The goal is to evaluate DeepSeek R1 within a specific application. This is not a comparison with other models or a study of whether R1 reasons better than DeepSeek V3.2. Nor does it assume that text generated between reasoning tags faithfully describes an internal process.
First, Pin Down the Version and Configuration You Are Testing
Before comparing prompts, record which model and inference path each run uses. The official model card on Hugging Face, the generation configuration file, and the repository’s chat template are useful references for documenting these details. In particular, the template file helps establish how messages are serialized. Two conditions that appear different in an interface may be converted into different token sequences, or two tests may cease to be comparable if the template changes between them.
Record the exact weights identifier or revision, the runtime and its version, the chat template, generation parameters, and any additional application mechanisms such as tools, validators, retries, or post-processing transformations. Do not change the template, temperature, and runtime at the same time. If several variables change together, the result will not clearly show which factor caused an improvement or deterioration.
If you use a compatible service rather than running the weights in your own runtime, document which parameters it accepts and which are outside your control. A compatible interface does not, by itself, guarantee identical serialization or identical inference details. If you cannot verify a detail, record it as an uncertainty in the report instead of presenting it as a known condition.
Minimum preparation before running tests
- 01Identify the weights revision and the specific DeepSeek R1 variant.
- 02Record the runtime, chat template, and generation parameters.
- 03Save the exact inputs, instructions, available tools, and expected output schema.
- 04Before testing, define what counts as a correct answer, an error, an abstention, and an integration failure.
Turn the Recommendations into Comparable Factors
Design conditions so that each comparison answers a defined question. To study instruction placement, keep the instruction content equivalent and compare, for example, a condition that puts instructions in the system message with one that puts them in the user message. If you also want to compare different positions within the user message, create a separate condition and keep the instruction text unchanged. This avoids attributing an effect to placement when it could instead be due to a change in wording.
Evaluate the “<think>” prefix as a separate factor. Compare a condition in which the assistant’s response begins without a forced prefix with one that adds the prefix specified in the documentation, checking how it is serialized in the actual prompt. Do not combine this in the same comparison with a change to the system message: if a condition changes both, an observed difference will not reveal which change caused it.
The recommendation to use a prefix does not show that it improves every task or is necessary in every runtime. Nor does it support the conclusion that visible reasoning text directly measures how the model operates internally. Evaluate what the application receives and can verify: the final answer, structure, latency, errors, and consistency.
A simple matrix of test conditions
| Factor | Condition A | Condition B | What must remain constant |
|---|---|---|---|
| Instruction placement | Instruction in the system message | Equivalent instruction in the user message | Task, wording, model, template, and parameters |
| Position within the user message | Instruction at the beginning | Instruction in another defined position | Instruction content and the rest of the prompt |
| Response prefix | No forced prefix | With the “<think>” prefix as specified by the tested serialization | Instruction placement and generation configuration |
Include Tasks That Represent Different Output Contracts
A test set made up only of mathematical problems may not represent an application that also extracts data, answers short questions, and processes documents. Build a small but varied set based on tasks the product actually performs. Variety is not just for show: it can reveal that a configuration helps one kind of task while harming another.
For a short factual answer, define in advance what information must appear and what additional content would count as a violation. For structured extraction, specify the schema, required fields, and how missing data should be handled. For a verifiable mathematical task, keep the correct answer and the method for checking it. For a document-based task, provide the same evidence to every condition and evaluate whether the answer is supported by it, distinguishes what is not documented, and avoids adding unrelated information.
Also include cases where the correct behavior is to abstain or state that the available evidence is insufficient. Without such cases, an evaluation may reward confident but unsupported answers. Keep examples of known errors and edge cases, but separate data used to tune a template from data used for the final evaluation. If instructions are revised during the experiment, record the change and rerun comparable conditions.
Measure Quality and Compatibility Separately
The primary score should reflect whether the response solves the task correctly according to criteria written down before looking at the results. Do not reduce everything to an overall impression. Also record instruction following, format validity, missing or unexpected fields, appropriate abstentions, and inappropriate abstentions. In applications with parsers, track whether the response can be processed without manual repair as an operational metric.
Measure latency consistently and define which interval you are measuring: generation, the full call, or the time observed by the user. Also record execution errors and retries. Repetitive or empty responses deserve their own category, with clear rules for identifying them; do not hide them inside an average quality score. An answer may contain a correct portion and still be unusable if it repeats text, omits a structural closing element, or breaks the required format.
Run each condition several times when the system introduces variation. Keep the individual results as well as any average or summary. Repetition can still help check environment stability for deterministic inputs and generation conditions that allow it; when there is randomness, report the number of runs and the observed spread. Do not present a small difference as a robust effect without showing how many cases support it.
Metrics worth reviewing together
| Dimension | What to record | Example of a regression signal |
|---|---|---|
| Correctness | Accuracy against a defined answer key or rubric | Fewer correct answers on a high-priority task |
| Instruction following | Compliance with constraints and requirements | Adds an explanation when a brief answer was requested |
| Format and integration | Structural validity and parser failures | Invalid JSON or omitted required fields |
| Abstentions | Appropriate and inappropriate abstentions, recorded separately | Answers without evidence or refuses despite sufficient evidence |
| Stability and performance | Repetition, empty outputs, latency, and variation | More repetitive outputs or longer response times |
Interpret Mixed Results Without Hiding Regressions
Suppose one condition improves the score on mathematical problems but increases formatting failures in extraction. It would be misleading to describe it simply as “better.” The decision depends on which tasks are critical, how costly each failure is, and whether the application has reliable mechanisms to detect and recover from it. An average improvement can conceal a serious regression in a particular function.
First compare each task separately, then present an overall summary and explain the aggregation method. If you use weights, state who chose them and what relative importance they represent. For small test sets, include absolute counts alongside percentages: moving from zero to one failure and from ten to eleven failures do not describe the same situation, even if an aggregate rate looks similar. Do not change the weights after seeing which condition comes out ahead.
DeepSeek’s documentation recommends multiple tests during evaluation. In practice, that means retaining variation across runs rather than selecting only a favorable answer. It also means not generalizing beyond the tested set: a result from a group of application-specific tasks is evidence for that application and those conditions, not a universal guarantee about R1.
Do Not Confuse Reasoning Text with a Reliable Explanation
The “<think>” label is part of the text interface of certain R1 conversation formats, but seeing text under that label does not show that you are viewing a complete or literal transcript of an internal process. Evaluation should focus on observable, verifiable outcomes. For a mathematical answer, check the answer; for an extraction, compare the fields with the document; for an abstention, confirm whether the available evidence was insufficient.
The paper published as DeepSeek-R1 Thoughtology analyzes characteristics of R1’s reasoning behavior and discusses topics such as length, controllability, and rumination. That context helps justify tracking repetitive outputs and length, but it does not replace application-level measurements or establish in advance what will happen when a template changes. A long explanation is not, by itself, evidence of a better answer; a short explanation does not prove that the task was solved without reasoning.
If the product must not display or store this content, separately check what the runtime exposes, what the application stores, and what the user receives. Do not assume that hiding a string in the interface also removes it from logs or traces. These properties depend on the integration and require their own checks.
Document the Decision as a Regression Report
A useful report makes it possible to repeat the experiment and understand why a condition was chosen. Include the exact model and template revisions, generation configuration, full prompts, task set, and scoring rules. Summarize results by task type, not just with one aggregate figure. Keep representative examples of successes and failures, handling sensitive data according to the team’s policies.
Describe what changed between conditions and what remained constant. Also note limitations: for example, whether the service exposes the weights revision or whether certain parameters cannot be controlled. If the result is inconclusive, that is a valid conclusion; you can expand the task set, repeat the test, or keep the current configuration while investigating. Do not turn a lack of evidence of regression into evidence that no regression exists.
Make the final operational recommendation specific: which template is accepted, for which tasks, with which version, and under what thresholds. If a variant works well only for one workflow, limit it to that workflow rather than presenting it as a universal template for DeepSeek R1.
A short report template
- 01Objective and tasks covered; success and rejection criteria.
- 02Model revision, runtime, chat template, and parameters.
- 03Compared conditions and the single intended difference in each pair.
- 04Individual and aggregated results by task, including failures and latency.
- 05Regression examples, known uncertainties, and limits on generalization.
- 06Decision: accept, reject, restrict to one workflow, or repeat the evaluation.
Practical Criterion: Keep What Passes the Application’s Test
The most useful prompting approach for a DeepSeek R1 application is the one that improves its tasks without weakening its output contracts or its safety and recovery mechanisms. To find it, establish a baseline, isolate changes, include varied tasks, repeat tests where appropriate, and show disaggregated results. A prefix, the position of an instruction, or the absence of a system message should not be accepted out of habit or rejected on intuition: test their effects under recorded conditions.
When results are consistent and meet the defined thresholds, deploy the configuration while monitoring the metrics that informed the decision. If the environment, weights, template, or tasks change, rerun the relevant evaluations. This way, official recommendations remain useful as initial hypotheses, while evidence from the application determines the final configuration.
Open questions
- Results depend on the exact weights revision, runtime, chat template, and parameters used; there is no universal experimental result for the proposed conditions.
- Parameter availability and control may differ between running your own weights and using a compatible service.
- The effects of the “<think>” prefix, instruction placement, and system message must be measured using each team’s specific tasks and conditions.
- The cited study of reasoning behavior provides context about length and rumination but does not demonstrate the causal effect of the protocol’s template variations.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction