Ilustración editorial para Un preprint alerta de que la destilación con información privilegiada puede generar agentes demasiado seguros
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

The central question: what does an agent learn when its teacher knows more?

A preprint titled *From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents* examines a training problem for agents that complete tasks over multiple steps or turns. Its thesis is that one form of self-distillation can teach a student agent to respond confidently without having, during the interaction, the information that justified the teacher’s response. The authors describe a potentially important consequence: in some reported results, the trained agent performed worse than ordinary reinforcement learning and, in the worst case reported, even worse than the untrained base model.

Here, “privileged” does not necessarily mean that the information is secret or always inaccessible. It refers to information added to the teacher’s view during training but not given to the student when it acts normally. The paper’s abstract does not specify every form of this information in each experiment. It does, however, describe a short, task-specific instruction written by an analyzer model as the kind of guidance that the proposed method keeps available during practice.

In the on-policy self-distillation setup described by the authors, the teacher and student start from the same model, but the teacher receives a view enriched with privileged information. The student then receives token-level supervision based on that view, despite not seeing the additional information when carrying out the task. The concern is an information asymmetry: the teacher can produce an answer consistent with a clue the student does not have, while the student learns to imitate the outcome without acquiring the clue.

The study does not show that all distillation produces this effect or that all agents will behave this way. Its claim is bounded by the configurations, models, and tasks evaluated in the preprint. That distinction matters: the paper presents an experimental finding, not a universal rule about agent systems.

02

What the preprint says about behavior and performance

The paper’s abstract characterizes the failure as confident behavior without the information that supports it. According to the authors, the trained student behaves as if it had privileged information it never observed. In multi-turn tasks, that difference may matter more than it would in a one-off question: the agent makes a succession of decisions, and an unjustified assumption early on can affect later actions. This is an interpretation of the problem described, not an additional quantitative result from the abstract.

The preprint reports experiments on AppWorld and SWE-bench Verified with three different student models. It says that the performance of the self-distillation method can fall well below plain reinforcement learning and, in the worst case, below the untrained base model. The supplied abstract does not identify which model or task produced that worst case, or specify its magnitude. Without the experimental details, that result cannot be attributed to a particular task.

The paper then presents Privileged Self-Practice (PSP), an alternative that changes where privileged information is used. Rather than putting it into the distillation loss signal, PSP keeps it in the practice context. When most of the student’s rollouts on a task fail, an analyzer model writes a short instruction for that task. The system samples the task again with the instruction in context and trains on the result using the unchanged GRPO objective.

The key distinction is not that PSP removes privileged information. It retains the information as guidance in the prompt during the new practice attempt and, according to the abstract, keeps it out of the loss. The agent can therefore generate a new trajectory under that guidance instead of directly learning to imitate a teacher distribution formed using information absent from the student’s view. The abstract does not explain exactly how the system determines that most attempts on a task have failed, or describe every implementation detail.

How the methods differ, according to the abstract

AspectSelf-distillation described in the paperPrivileged Self-Practice
Use of privileged informationConditions the teacher’s view during supervision.Retains it in the practice context as a task-specific instruction.
Where the signal entersThe student learns from supervision derived from the teacher’s enriched view.The information stays in the prompt and, according to the abstract, does not enter the loss.
Response to failuresThe abstract does not describe an equivalent retry mechanism here.When most rollouts for a task fail, an instruction is added and the task is sampled again.
Training objectiveOn-policy self-distillation.The authors say the GRPO objective is unchanged.
03

What comparisons the paper reports—and how to read them

On AppWorld and SWE-bench Verified, the preprint reports that PSP achieved the best average score in every evaluated setting and was the only method to consistently outperform plain GRPO. It also reports improvements of up to 65% in task-goal completion on AppWorld and up to 61% in the resolved rate on SWE-bench Verified. These are the maximum improvements cited in the abstract; they should not be read as gains achieved by every model on every test.

The abstract does not clarify whether those percentages are relative increases or percentage-point changes, and it does not provide absolute scores, uncertainty intervals, or results broken down by model. Nor can the abstract alone establish whether the reinforcement-learning comparison exactly controls for total training budget, data, or the number of attempts. Those details are necessary to assess how much of the difference is attributable to the method rather than other experimental factors.

There is also a distinction in how the comparisons are described that should be preserved. The abstract says that the worst case fell below the base model and that performance fell below plain reinforcement learning; it later describes PSP’s comparison with plain GRPO. Without the methodological details, we should not assume these labels refer to exactly the same comparator in every evaluation. PSP’s favorable results are the authors’ report, not an independent confirmation.

The PSP cycle, as described in the available material

  1. 01The agent attempts a task and generates practice rollouts.
  2. 02If most of those rollouts fail, an analyzer model writes a short instruction tailored to the task.
  3. 03The task is sampled again with that instruction included in the context.
  4. 04The result is used for training with the unchanged GRPO objective; the privileged information remains in the prompt and is not incorporated into the loss.
04

What remains unresolved before generalizing

The scope of the result is limited by the methodological information in the abstract. It names the two benchmarks, says that three student models were evaluated, and gives an outline of the PSP mechanism. It does not specify the models’ names and sizes, the exact composition of the data, the number of runs, the full success criteria, or the threshold that triggers guidance. Without these details, readers cannot reproduce the comparison or fully assess how sensitive it is to configuration choices.

Another open question is whether PSP consistently improves behavior when the agent does not have access to privileged information. The described mechanism suggests that the instruction is provided during practice on the selected task, but the abstract does not offer a detailed test of generalization to situations where that instruction is unavailable. The reported result should therefore not be turned into a claim that the method fully solves the problem of unjustified confidence.

The preprint proposes an explanation and an experimental alternative; the available material does not establish that it has passed peer review or provide an independent evaluation. Related work examines other problems in privileged-information distillation, but does not by itself validate PSP’s results. The cautious takeaway is specific: the study argues that training from a more informed teacher view can create mismatches in some multi-turn agents; it reports a promising alternative in two environments; and it leaves open details that would determine how robust and broad the advantage is.

In practical terms, it is worth checking not only whether an agent reaches the correct answer but also what information it actually had when making its decisions. When evaluating a training technique of this kind, it helps to distinguish the information available to the teacher, the information available to the student during practice, and the information available at evaluation time. That separation can reveal whether apparently capable behavior depends on a clue the agent will not be able to consult after deployment.

05

Explore more on Inferama

For broader coverage of these developments, explore Inferama’s News, Compare, and Discover sections. Together, they cover AI updates, side-by-side comparisons, and solutions readers can explore.

These sections serve different editorial purposes: News follows current developments, Compare examines tools side by side, and Discover highlights solutions to explore. This article does not recommend a particular product.

Open questions

  • The abstract does not specify exactly what privileged information the teacher and student received in each experiment, beyond describing short, task-specific instructions generated by an analyzer for PSP.
  • It does not identify which task or model produced performance below the base model, or provide the magnitude of that decline.
  • The abstract does not establish whether comparisons with reinforcement learning control for training budget, data, number of attempts, or other factors.
  • It does not say whether the maximum reported improvements are relative percentages or percentage-point changes.
  • There is not enough evidence to claim that PSP improves behavior when privileged information is unavailable during evaluation.
  • The supplied primary source is a preprint; the available material does not establish peer review or independent validation.
06

Keep exploring

06

Sources consulted

SECONDARYEviSD: Auto-destilación condicionada a la evidencia para agentes potenciados por búsqueda | alphaXivwww.alphaxiv.org · 25 Sep 2026↗ SECONDARYMás allá de la imitación absoluta: Guía residual anclada para destilación privilegiada en política | alphaXivwww.alphaxiv.org · 25 Sep 2026↗ SECONDARYPrivilegiados, pero sesgados: Cómo los profesores condicionados por PI rompen la autodestilación | alphaXivwww.alphaxiv.org · 25 Sep 2026↗ SECONDARYDOPD: Destilación Dual de Política | alphaXivwww.alphaxiv.org · 25 Sep 2026↗ PAPER · PrimaryDOPD: Dual On-policy DistillationarXiv · 25 Sep 2026↗ PAPER · PrimaryPrivileged, but Biased: How PI-Conditioned Teachers Break Self-DistillationarXiv · 25 Sep 2026↗ SECONDARYUn Síntoma, Tres Palancas: Una Revisión Crítica de la Autodestilación Basada en la Política | alphaXivwww.alphaxiv.org · 25 Sep 2026↗ SECONDARYCuando la Guía Privilegiada se Desalinea: Enrutamiento Alineado con el Estado y Auto-Destilación Contextualizada para Agentes Multiturno | alphaXivwww.alphaxiv.org · 25 Sep 2026↗ SECONDARY¿Qué aporta la información privilegiada a la autodestilación basada en políticas? | alphaXivwww.alphaxiv.org · 25 Sep 2026↗ SECONDARYAnti-Autodestilación para RL de Razonamiento mediante Información Mutua Puntual | alphaXivwww.alphaxiv.org · 25 Sep 2026↗ SECONDARYSobre docentes repulsivos y atractivos: separar la corrección del comportamiento en la autodestilación | alphaXivwww.alphaxiv.org · 25 Sep 2026↗ PAPER · PrimaryFrom Self-Distillation to Self-Practice: Privileged Information for Multi-Turn AgentsarXiv cs.AI · 25 Sep 2026↗ PAPER · PrimaryWhat Does Privileged Information Add to On-Policy Self-Distillation?arXiv · 25 Sep 2026↗
03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction