Ilustración editorial para Mantener repositorios con IA: qué delegar y qué revisar
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

Maintaining a Repository Is Not the Same as Requesting a Change

Repository maintenance involves distinct activities: understanding existing code, diagnosing failures, updating documentation, fixing defects, changing dependencies, and checking that a modification does not harm other behavior. Each activity calls for different information and safeguards. That is why “using AI for maintenance” does not describe a single level of delegation. It could mean asking for an explanation, requesting a proposed change, or allowing an agent to edit files and run tools.

The important distinction is between producing a suggestion and taking responsibility for a task with a verifiable outcome. A coherent description of a module does not prove that its behavior has been understood; a plausible-looking patch does not prove that it fixes the defect; and a passing test provides evidence only about what that test checks. The ability to edit files or run commands does not, by itself, show that the system chose the right scope or that the change is safe to integrate.

The practical thesis of this guide is that the appropriate degree of autonomy should depend on the task, the potential impact of getting it wrong, and the evidence the team can inspect. For exploratory tasks, AI can help generate hypotheses or summarize code. Changes that affect data, public interfaces, permissions, dependencies, or deployments call for stricter limits and explicit approval. The decision is not “AI or no AI,” but what it may do, under what conditions, and who is responsible for accepting the result.

It is also worth distinguishing a tool’s performance in a demonstration from its suitability for a particular repository. A demonstration does not necessarily reveal how much context the tool received, which tools were enabled, what checks were run, or how much review the result required. The team’s own experience with representative tasks is a more useful basis for a decision than a general impression of what a model can do.

02

Classify the Task Before Assigning Autonomy

A simple classification starts with the type of result requested. Understanding and documenting usually produces explanations that a person can check against the code. Diagnosis produces hypotheses about a cause and requires distinguishing them from observed facts. Proposing a change produces a diff that must be reviewed. Executing changes and tests adds the possibility of affecting files or processes. Changing dependencies or interfaces can affect consumers that are not visible in the immediate context of the task.

The available context affects the difficulty. An issue with reproducible steps, relevant logs, and a failing test provides a more concrete signal than a vague report. A narrowly scoped documentation update may be verifiable if the team knows the current behavior; describing an obsolete function first requires establishing the source of truth. If the request does not define what “done” means, AI may produce a polished response without meeting the actual need.

The table is not a universal rating of tasks. It summarizes a starting guideline that should be adjusted to the repository and the consequences involved. “Reversibility” here means how easy it is to detect and undo the change, not that every error is harmless. Even a narrowly scoped modification can have a high impact if it affects authentication, persistent data, or a public contract.

An Initial Matrix for Choosing the Level of Involvement

Use the matrix as a starting point. If a task fits more than one row, apply the stricter level of control until the team can justify another approach.

Task typeTypical risk if wrongCondition for proceedingRecommended involvement
Explain a module or summarize documentationLow to medium: an incorrect explanation can misdirect later workSpecific references to files, symbols, or documentation that can be checkedAssistance; a person validates the facts before relying on them
Investigate a failure and propose a causeMedium: an incorrect hypothesis can derail the diagnosisHypotheses separated from observations, with reproducible evidenceUse AI to investigate; confirm the diagnosis with tests or review
Edit documentation or fix a contained defectVariable: depends on scope and the consequences of the behaviorSmall diff, clear expected outcome, and relevant checksAI may prepare the change; a person reviews the diff and the result
Update dependencies or change a public interfaceMedium to high: compatibility, security, or external consumers may be affectedAn update plan, identified impact, tests, and approval by an accountable reviewerKeep the work scoped and require human review before integration
Change security controls, data, or deploymentsHigh: the effects may extend beyond the local changeAuthorized scope, specialist assessment, and independent controlsAI may support analysis or prepare a proposal; it must not approve or deploy on its own
03

Decide Using Five Factors, Not a Label

For a specific decision, assess five factors: impact of an error; reversibility; test coverage and relevance; code sensitivity; and adequacy of the context. You do not need to turn them into a mathematical score. An average can obscure the fact that a task has one critical factor, such as changing permissions or having no reliable way to verify the result.

Impact asks what could happen if the change were wrong: anything from confusion caused by incorrect documentation to data loss or a service interruption. Reversibility asks whether the change can be detected and undone before it causes lasting consequences. Test coverage considers whether checks exist for the affected behavior, not merely whether the repository has a test suite. Sensitivity identifies areas that call for particular expertise or authorization. Context includes requirements, versions, project conventions, and relevant dependencies.

A conservative rule is to increase supervision as impact rises, reversibility falls, or tests are missing. If permissions, affected consumers, or external effects are unknown, do not treat that lack of knowledge as low risk: limit the work to investigation and request more information. “Stop” is a valid option when a safe check cannot be defined or the system cannot respect the requested scope.

These categories are a decision framework, not a claim that every change in a category carries identical risk. For example, a minor update may be routine in one project and critical in another if it affects a central library or an exposed component. The repository owner must provide that context and review the classification.

Decision Path Before Starting a Task

If an answer is unknown, do not assume the requirement is satisfied: narrow the scope or request a review.

  1. 01Define the expected outcome in observable terms: affected files, behavior that must be preserved, and the condition for completion.
  2. 02Identify whether the work is exploratory, documentary, diagnostic, a modification, or an operation with effects outside the repository.
  3. 03Assess impact, reversibility, available tests, code sensitivity, and the context provided.
  4. 04Set permissions and scope limits before allowing edits or commands.
  5. 05Require evidence proportionate to the risk: references, hypotheses, diff, tests, and limitations.
  6. 06Assign a person to review and approve changes that may affect behavior, dependencies, interfaces, or controls.
  7. 07If the result cannot be verified or the effects cannot be contained, stop execution and hand off the task.
04

Limit the Work Before Allowing Changes

Controls should be defined before the system starts acting, not after discovering that it ran something unexpected. For an initial trial, use an isolated working branch, specify which files or components are in scope, and list actions that are not allowed. Restrict access to credentials and sensitive data; an ordinary maintenance task normally does not need publishing, deployment, or secret access permissions.

Distinguish between reading, proposing, editing, and executing. You can allow AI to inspect files and suggest commands without authorizing it to run them. If execution is enabled, define which commands are acceptable and which require approval. Check whether a tool can modify files unrelated to the change, install packages, access the network, or perform actions with side effects. Specific capabilities and controls depend on the product and its configuration; do not infer their limits from the interface or marketing description.

Minimum necessary permissions reduce the possible damage, but they do not replace review. A separate branch does not prevent a diff from containing an incorrect modification; a test environment does not prove that there are no effects on external services. Similarly, allowing only certain commands does not guarantee that the results of those commands are sufficient. The team must check which permissions are actually active and observe the actions taken.

For changes to dependencies, public interfaces, security controls, or deployment processes, require approval before integrating or executing in a shared environment. AI can gather information, prepare a proposal, and run authorized checks; accepting the risk remains a team decision. Do not grant publishing or deployment permissions solely to save review steps.

05

Require Evidence That Can Be Checked

A useful response should let another person reconstruct what was done and why. For an explanation, request references to files, symbols, or documentation that support its claims. For a diagnosis, separate observations from hypotheses: “the test fails in this case” is different from “this line causes the failure.” For a change, require a readable diff, a description of scope, the checks performed, and known limitations.

References should be specific and verifiable in the repository. A list of filenames is not enough if the reviewer cannot connect them to the affected behavior. When tests are reported, specify which ones were run, under what conditions, and whether they completed successfully. A generic statement such as “all tests pass” is insufficient if it is unclear which test set was run or whether relevant tests were omitted.

Evidence does not automatically make a change correct. A test suite may not cover an edge case; an explanation may cite the right file but misinterpret the logic; a small diff may break an interface used by another component. Review should compare the evidence with the expected outcome and look for effects that the checks do not cover.

Also document what was not checked. If integration tests were not run, a required service was unavailable, or the agent could not access a submodule, those limitations should appear in the handoff. Explicit uncertainty lets the team choose whether to complete a check, knowingly accept a limitation, or stop the change.

06

Testing, Review, and Approval Are Different Controls

Automated tests verify conditions defined by the project. They may be unit, integration, type, formatting, or domain-specific tests, but none of those labels guarantees that they cover the change. Before delegating a modification, determine which tests relate to the affected behavior and what additional check might be needed. For a documentation task, for example, validation may include checking that the instructions can be followed and match current behavior.

Diff review looks for things an automated result may miss: out-of-scope changes, unjustified assumptions, error handling, compatibility, data exposure, and effects on consumers. Approval is an integration decision assigned to a person or team policy; it is not synonymous with a test passing. In repositories with protected branches, a team can configure review and check requirements before allowing integration. These mechanisms support a control process, but do not by themselves establish whether the review was substantive.

Whenever possible, define acceptance criteria before making the change. Include the behavior that must be achieved, behavior that must remain unchanged, required checks, and who can approve. For sensitive areas, add review by someone with the relevant expertise. If no test covers the main risk, consider creating one before or alongside the change; do not substitute a convincing agent explanation for that gap.

Do not confuse local success with operational safety either. A command may complete successfully in the working environment without reflecting the production configuration. A change may compile and still alter an interface. For operations that affect external systems, define specific checks and approvals outside the repository, and avoid making execution implicitly authorized just because editing was delegated.

Acceptance Gate for Integration

Adjust these controls to the impact of the change; a high-risk condition is not offset by favorable answers to the others.

  1. 01The change meets a defined expected outcome and stays within the authorized scope.
  2. 02The diff has been reviewed and contains no unexplained or out-of-task modifications.
  3. 03Relevant tests were run; failures and omitted tests are explained.
  4. 04Compatibility, security, data, and external-effect risks were assessed as appropriate.
  5. 05The person with suitable authority and expertise approved the change.
  6. 06Permissions to integrate, publish, or deploy remain subject to their own controls.
07

Correct, Limit, or Stop the Intervention

Continue with assistance when the scope is clear, tools have appropriate permissions, and the handoff provides verifiable evidence. Request corrections when references are missing, the diff includes unrequested work, relevant tests were not run, or the explanation mixes facts with conjecture. Rather than simply saying “fix it,” state which condition is unmet and request a new proposal limited to that point.

Stop work if the system tries to access unauthorized resources, cannot respect the file boundary, proposes commands with effects the team cannot control, or cannot explain a high-impact change. It is also sensible to stop when the task depends on missing requirements, specialized knowledge that has not been provided, or behavior that cannot be tested safely. Stopping does not mean declaring the tool useless: it means that this task, with this context and these permissions, is not ready to be delegated.

If an unexpected change appears, preserve the action log and check the repository state before continuing. Do not automatically accept a second proposal just because it corrects the first error: verify scope, tests, and consequences again. For tasks affecting persistent data, security controls, or production, use the team’s usual procedures to assess and roll back changes rather than improvising a repair in the same session.

Operational Decision During Execution

The response should depend on the evidence observed, not on how confidently the explanation is phrased.

SignalActionCondition for resuming
The proposal is in scope and includes relevant testsContinue with the planned reviewThe diff and its effects remain acceptable
References are missing or a relevant test has not been runRequest evidence or an additional checkThe new handoff addresses the gap and states its limitations
Files or commands outside the authorized scope appearStop execution and review the changes madeThe team restores clear limits and confirms the repository state
Risk is high and no adequate verification existsDo not integrate; refer for specialist review or design a testAn accepted path for validation and approval exists
08

Run a Pilot Using Your Team’s Own Tasks

Before expanding use, select a small set of historical or reproducible tasks from the repository. Include variety: explain a module, investigate a known failure, update a documentation section, and prepare a contained modification. Define in advance what counts as a correct answer, which tests are relevant, and what work a person must do. Do not choose only easy tasks that are likely to create a favorable impression.

For each task, record whether the result was accepted, corrected, or rejected; what defects it introduced or failed to detect; which tests it ran; how much team time review took; and how much work had to be redone. Also record the execution conditions, such as the context provided, enabled permissions, and environment limitations. Without this information, comparisons across attempts can confuse configuration differences with quality differences.

Interpret metrics alongside reviewed examples. A shorter time to get a diff does not necessarily mean an improvement if review effort or the number of corrections increases. An isolated acceptance rate may hide the fact that the selected tasks were unrepresentative. Decide what results would justify expanding, maintaining, or reducing the pilot, and who is authorized to make that decision.

The pilot should test the complete workflow, not just the ability to generate changes: permission limits, traceability, tests, review, and approval. Include at least some cases where the right response is to ask for clarification or stop. If the process only measures how many tasks are completed, it may encourage execution even when context or verification is missing. The goal is to determine where assistance is useful and under what controls.

09

Final Checklist for Choosing a Level of Use

A practical decision starts by naming the task and its expected outcome, not by asking what level of autonomy a tool offers. Then check whether the necessary context is available, whether permissions can be limited, and whether tests can detect the important errors. If any answer is no, narrow the task to investigation or a proposal and do not authorize integration.

Approval requirements should match the impact. For a low-risk documentation change, an ordinary review and an accuracy check may be enough. For dependencies, public interfaces, security controls, or deployment processes, require clear accountability and additional validation. Team policy should specify who can approve, which checks block integration, and which actions require separate authorization.

Tool documentation can clarify the permissions and controls a specific product offers, but it does not prove that they are enabled in a particular configuration or that a task’s result is correct. Verify the actual environment and record decisions. Research on code review involving people and agents can provide context about ways to collaborate, but it does not replace evaluating a workflow in the repository and team where it will be used.

In short, first delegate work that can be scoped and verified; request evidence another team member can inspect; retain human approval for changes whose impact warrants it; and stop execution if permissions cannot be controlled or the outcome cannot be demonstrated. Adjust the level of use based on pilot data, and revisit the limits when the repository, tools, or consequences of a task change.

Open questions

  • The risk level of a task depends on the repository, its consumers, the execution environment, and the specific consequences of an error.
  • Agent capabilities and permissions vary by product and configuration; they must be checked in the actual environment.
  • The existence of tests does not prove that they cover every case relevant to a change.
  • The supplied arXiv source studies code-review conversations involving people and agents, but the information provided does not support attributing specific quantitative findings to it or generalizing it to all repositories.
  • Tool documentation describes available controls but does not confirm which controls are enabled in a particular installation.
10

Keep exploring

10

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction