The decision: pay for a task completed, not an impression
Choosing between Gemini 3.1 Pro and Gemini 3.7 Flash for code maintenance is not simply a matter of asking which one is “better.” For an engineering team, the practical question is which model can complete a particular task with a correct patch, how much it costs to reach that result, and how long it takes. A model that responds quickly but introduces a regression or requires several rounds of correction may be less efficient than one with a more expensive call.
A useful comparison needs to measure completed work. That means checking whether the change addresses the issue, passes relevant tests, and avoids modifying parts of the codebase unrelated to the request. It also means accounting for failed attempts, tool use, retries, human intervention, and runs that do not end in an acceptable patch. Cost per call alone does not describe the cost of maintaining a repository.
No results are available here from a test run on the same repositories under a shared harness. It would therefore be unsound to claim that Flash is sufficient for routine work, that Pro justifies its cost on complex tasks, or that either model wins. What follows is a protocol for finding out, along with guidance for interpreting the results without presenting hypotheses as facts.
Which models are being compared—and what to verify before starting
The first step is to pin down the exact identity of each system. The official model page consulted identifies Gemini 3.1 Pro with the ID `gemini-3.1-pro-preview` and labels it a Preview version. The reasoning guide consulted includes `gemini-3.7-flash` and `gemini-3.1-pro-preview` among the models for which it provides configuration information. Do not replace these names with shortened labels in the test records: the ID sent to the API is part of the experimental setup.
Availability and status can change. Before running the test, and again when it closes, verify which IDs are accepted, which access channel is being used, and whether either model has become unavailable or changed status. If a model changes during the test, the article should state which runs correspond to each version or separate the periods; it should not silently treat them as one homogeneous sample.
Also record the provider and access channel, date, generation parameters, context limit, stopping rules, and any reasoning configuration. The official documentation describes options and defaults, but that does not prove that two settings with the same name represent the same amount of computation, internal budget, or behavior. A comparison using “equivalent” settings is defensible only if equivalence is operationally defined and the actual parameters used are reported.
Do not copy prices from an old table or infer them from the model name. Verify the applicable billing components on the date of the test and explain what is included: input, output, tool calls, or other relevant charges. If access goes through an intermediary or uses a rate different from direct API access, measure and describe that channel, because the results do not automatically transfer to another one.
Minimum pre-test record
Complete and publish these details before interpreting differences between the models.
| Item | What to record | Why it matters |
|---|---|---|
| Identity | Exact ID sent, provider, and access channel | Avoids attributing results to an ambiguous label or a different version. |
| Status | Availability and status at the beginning and end | Makes it possible to detect version changes or interruptions. |
| Configuration | Reasoning, generation, context, and limits | Makes the run reproducible without presuming that settings are equivalent across models. |
| Billing | Current rates and included components | Allows cost to be calculated per attempt and per accepted task. |
| Environment | Repositories, commits, tests, tools, and permissions | Defines what work each model was able to perform. |
Design: repositories, issues, and acceptance
Freeze the test set before running either model. Each task needs an identifiable repository and base commit, an issue description, a reproducible environment, and an acceptance criterion written in advance. Changing tasks or rules after learning which model produced each patch increases the risk of tuning the evaluation to the results.
Select issues from different repositories and group them by type and difficulty: diagnosing a bug, making a contained fix, changing several files, and handling a task where existing tests do not fully cover the expected behavior. Publish the distribution. A collection dominated by small changes may favor one work profile; a collection made up mainly of broad problems may favor another. Variety reduces that risk but does not eliminate it.
The issues should be genuine and relevant to maintenance, but selection requires care. The SWE-bench paper describes a precedent based on problems derived from real GitHub issues. Its dataset and harness documentation can inform the design of an evaluation, but using benchmark tasks does not guarantee that a particular test measures the work of a given team well. OpenAI has also published a warning about limitations in SWE-bench Verified. Attribute that warning to OpenAI’s stated position; do not present it as an independent audit of the models compared here.
To reduce contamination, check whether the solution or an equivalent patch is already present in the context provided, the task files, or materials available to the system. Also describe exactly what information each model receives: the issue, history, instructions, documentation, and tests. Saying that both received “the same prompt” is not enough if one had access to additional data or different tools.
Run tasks on clean copies of the same commit. Keep the harness, dependencies, and test commands identical. A reproducible environment helps separate the effect of the model from the effect of a machine, dependency version, or manual repository change. If infrastructure fails rather than the patch, label the run according to a rule set before seeing the results.
Evaluation sequence
Apply the same procedure to every task-and-model combination.
- 01Freeze the issue, base commit, tests, and acceptance criterion.
- 02Create a clean environment and give the model the same starting materials and permissions.
- 03Record calls, available billable token data, tools, timings, errors, and file changes.
- 04Apply the patch without manual intervention and run the predefined tests.
- 05Blindly review the relevance of the change, regressions, and unnecessary work.
- 06Classify the result using pre-established rules, and publish attempts that were not accepted as well.
Define “accepted task” before looking at the patches
Passing a test does not automatically make a solution acceptable. Existing tests may not cover the central requirement, and a patch may make them pass through an overly broad change or a modification that hides the bug. Acceptance should combine relevant automated tests with a structured human review.
At a minimum, the rubric should ask whether the patch resolves the issue, preserves unrelated behavior, introduces known regressions, changes only what is necessary, and remains maintainable. It must specify what happens when the result appears correct but the tests are insufficient. In those cases, independent checks can be added, or the task can be classified as uncertain instead of forcing a pass or fail unsupported by the evidence.
Reviewers should receive anonymized patches and apply the same rubric without knowing which model produced them. Record disagreements and establish a resolution procedure, such as a second review. Human evaluation is not infallible either, so publish the acceptance definition, rejection categories, and share of disputed decisions.
Two different views: equal budget and equal deadline
A comparison under the same budget and one under the same time limit answer different questions. They should not be collapsed into a single figure. For the first, assign each model a spending cap per task, including the billing components defined for the test. Measure how many tasks it completes with an acceptable patch before reaching that cap. A failed attempt uses part of the budget and must remain in the results.
For the second, give both models the same maximum time, measured from the start of the task to delivery. The clock should include waiting, tool calls, retries, and any other latency experienced by the user. If human review is measured separately, say so; if it is included, apply the same procedure to both. Otherwise, comparing only model response time with the total time of a workflow mixes different metrics.
For a fair comparison, set the spending cap, deadline, retry count, step limit, permissions, and stopping rule before running the test. If a model reaches the cap without producing an acceptable patch, the result is an unresolved task under that condition, not a run to remove from the analysis. This approach answers a concrete purchasing question: what can be achieved with a fixed amount of money or within a defined time window?
How to interpret the two limits
The same run can perform differently depending on which constraint matters most.
| Condition | Primary metric | Question answered |
|---|---|---|
| Equal budget | Accepted tasks within a spending limit and cost per acceptance | What share of useful work can be obtained with a fixed budget? |
| Equal time | Tasks accepted before the deadline and total latency | Which model delivers more acceptable work within a fixed time window? |
| No shared limit | Does not allow direct attribution of efficiency | May describe real-world use, but does not isolate the effect of the constraint. |
What to measure per task and per group
The central metric should be the acceptance rate: the proportion of runs that produce a patch accepted under the rubric. To make it interpretable, report it alongside the total number of tasks and attempts, not just as a percentage. A high rate across a small number of runs carries different uncertainty from the same rate across a larger set.
Report regressions, relevant tests passed, unnecessary changes, retries, service errors, human intervention, and tasks that end without a patch. Spending can be expressed per attempt and per accepted patch, provided the formula and billing components are stated. If a model does not resolve a task, its cost must not disappear from the denominator: excluding failures would make the cost per success look artificially favorable.
At a minimum, separate latency into model time, waiting or queue time when measurable, tool operations, and total time to delivery. For teams, review time and manual correction may also matter. Two models with similar generation times can impose different workloads if one requires more checks or repairs.
Break results down by task type and difficulty as well as providing an overall summary. An aggregate figure may hide that one model handles small changes better while the other avoids failures in changes spanning multiple files. Do not call a result “better” just because it leads on an average if the reader’s team works on a different task profile.
Repeated runs, service failures, and uncertainty
A model’s output can vary between runs, even when the task and environment remain unchanged. One run per issue is therefore not enough to describe stability. Set the number of repetitions before starting, justify it in light of the scope of the test, and keep it the same for each model and task. When service conditions prevent a run from completing, distinguish an infrastructure error from a model failure without deleting the data.
Show variation and uncertainty, not just an average. Counts, appropriate intervals, and per-task results can be published; the specific statistical method depends on the design and should be described. If the set is small, keep the conclusion narrow: a difference observed in those cases does not demonstrate that it will hold for other languages, repositories, teams, or levels of difficulty.
Sensitivity to the harness also matters. Changing the prompt, tool limit, context budget, or stopping rules can change the result. A controlled test identifies the effect of a particular configuration; it does not measure a universal capability independent of how the model is used. If additional tests use other configurations, present them as such rather than combining them with the primary result.
A decision matrix for engineering teams
Without test results, the matrix cannot assign an empirical advantage to Gemini 3.1 Pro or Gemini 3.7 Flash. It can, however, help translate collected data into a decision that fits the team’s work. The comparison must account for the minimum quality threshold: if a model does not meet the required acceptance rate or safety level, a low cost alone does not make up for it.
For routine, well-defined tasks, a team can first evaluate whether Flash meets that threshold within its budget and deadline. That is a hypothesis to test, not a property established here. For difficult tasks, extensive changes, or work with significant consequences, check whether Pro increases acceptance or reduces human intervention enough to justify the extra expense, without assuming that the Pro name guarantees that outcome.
In either case, retain human review when the risk of a change calls for it. If both models produce patches that often need correction, or the available tests cannot verify behavior, the reasonable conclusion may be that neither meets the automation criteria for that class of work. The decision may also be a hybrid one, provided the routing between models is measured rather than assuming that choosing by difficulty improves the result.
Practical rules once the data is available
These rules depend on verified results from the team’s own tasks.
| Finding in the test | Decision to consider | Caution |
|---|---|---|
| Flash meets the acceptance threshold and budget for contained tasks | Trial Flash for that task group with oversight appropriate to the risk | Do not extrapolate to complex changes or unevaluated repositories. |
| Pro improves acceptance or reduces corrections on complex tasks | Calculate whether the improvement justifies additional cost and latency | Compare total cost per accepted patch, not only the price per call. |
| Both fail frequently or cause regressions | Keep work manual or redesign the workflow and tests | Do not lower the acceptance standard just to produce a winner. |
| Differences are small or highly variable | Expand the sample or decide based on operational constraints | Do not turn an uncertain difference into a general conclusion. |
Materials for reproduction and limits of the conclusion
A comparison intended to guide an engineering decision should publish the protocol, included tasks, base commits, configuration for each model, limits, and acceptance rubric. Where permissions and licenses allow, it should also provide patches, anonymized logs, test results, and reasons for rejection. Exclusions and failed attempts are part of the evidence, not secondary details.
The harness can draw on reproducible evaluation practices, such as running patches in isolated environments and controlling how changes are applied. The SWE-bench harness documentation describes an approach using Docker environments to run and evaluate tasks. This is a methodological reference; using that harness does not, by itself, guarantee that a test represents a company’s workflow.
Results support conclusions only about the documented models, versions, tasks, configuration, and date. They do not automatically establish how other Google products, other versions, or other providers behave. Nor do they replace evaluation on private repositories with a team’s own security and review requirements. The question “When does paying more per resolved task pay off?” can be answered only after defining what counts as resolved and measuring the full cost in the relevant context.
For now, the rigorous editorial answer is not a winner, but a decision condition: compare both models using frozen tasks, blind acceptance review, shared spending and time limits, and published uncertainty. If those data show consistent differences that matter to the team’s work profile, an option can be recommended for that use. If not, the honest conclusion is that the available evidence does not distinguish them.
Publication checklist
Before making a recommendation, check that the report records each of these elements.
- 01IDs, statuses, access channels, and execution dates for both models.
- 02Reasoning and generation parameters, without declaring identically named settings equivalent.
- 03Tasks, repositories, commits, selection criteria, and contamination checks.
- 04Tools, permissions, limits, retries, and stopping rules.
- 05Acceptance definition, blind review, disagreements, and regressions.
- 06Results by task and category, including failures, missing data, spending, latency, and uncertainty.
- 07Rates verified for the relevant date and the method used to calculate cost per accepted patch.
- 08Exclusions, reproducible materials, and explicit limits on generalization.
Open questions
- No experimental results, run counts, evaluated tasks, acceptance rates, regressions, latencies, or costs are provided for these two models.
- Accepted IDs, availability, Preview status, and configurations can change; they should be verified on the execution dates and at editorial close.
- Current rates and billing data are not provided, so comparative cost cannot be calculated.
- Repositories, tasks, rubric, tools, limits, repetitions, and statistical method are unspecified; empirical recommendations depend on running that test.
- The warning about SWE-bench Verified comes from an OpenAI publication and should be attributed as OpenAI’s position, not described as an independent evaluation.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction