Ilustración editorial para ASIRF plantea adaptar la redacción de datos sensibles al contexto sin reentrenar el modelo
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

The problem: what counts as sensitive depends on context

A piece of data is not always sensitive in isolation. Its significance can depend on the domain in which it appears and the purpose for which it is requested or used. A name, a location, or a reference to an employment relationship, for example, might call for different treatment depending on the content and purpose. These examples are illustrative; the information available about ASIRF does not describe these specific cases.

The preprint starts from a limitation it attributes to some privacy filters and named-entity recognizers: their categories are fixed during training. Applying such a system to a new domain may therefore require retraining it. ASIRF proposes a different approach: while processing an input, it retrieves definitions of sensitive information associated with the relevant domain from a flexible knowledge base.

The proposal changes where some of the adaptation happens. Instead of relying only on categories learned in advance, the system consults definitions available at inference time. That does not, by itself, prove that it identifies every sensitive item, prevents disclosures, or is suitable for every domain. It describes the mechanism evaluated by the work, not a guarantee of privacy.

02

How ASIRF works according to the preprint

The available description presents ASIRF as an agent-based framework. Given an input, the system retrieves domain-specific definitions from a flexible knowledge base and uses them at inference time. The authors argue that this adaptation does not require retraining the model for each new domain.

The paper compares two variants: a multi-agent architecture that uses a three-call pipeline and a single-agent version. The information provided does not specify what each call contains, the agents’ instructions, how the domain is selected, or the exact rules for turning a detection into a redaction. It is therefore not possible to reconstruct the complete technical workflow rigorously from the available summary.

The paper also says that it uses a few dozen expert-authored definitions for each domain and does not require training data. This should be understood within the evaluated design: it does not mean that the system needs no definitions, configuration, testing, or human oversight. Nor does it establish that a small knowledge base covers every way sensitive information can be expressed.

General workflow described by the paper

  1. 01Receive the text to be analyzed.
  2. 02Retrieve definitions of sensitive information associated with the input’s domain.
  3. 03Use the agent system to decide what information to redact.
  4. 04Evaluate the output against the datasets and experimental baseline.
03

What the evaluation compares

The preprint summary reports an evaluation involving ten small open-weight models and eight datasets. The datasets include fictional out-of-distribution domains—that is, examples the paper presents as outside the training or reference domains. The description does not list all the models, datasets, or domains, or the criteria used to decide that a domain is new.

The main comparison mentioned is with OpenAI Privacy Filter, described in the summary as a trained-classifier baseline. The highlighted result is that ASIRF achieves higher recall than this baseline in 68 of 80 model-domain combinations, with at least one of the two architectures. According to the summary, the shortfalls occur mostly in domains from the baseline’s training distribution.

This aggregate figure is not enough to establish that the two architectures are equivalent or that ASIRF wins on every model, domain, or dataset. Nor does it reveal the size of the differences by itself. The summary provided does not give complete results by dataset, precision figures, over-redaction measurements, or detailed comparisons between the multi-agent and single-agent variants.

What can—and cannot—be concluded from the summarized data

Reported findingCautious interpretationWhat is missing to interpret it
Ten models and eight datasetsThe system was evaluated in more than one configuration and on domains described as out of distribution.The complete list, disaggregated results, and the operational definition of an unseen domain.
Higher recall than OPF in 68 of 80 combinations, with at least one architectureThe summarized result favors ASIRF’s recall in most of the combinations.Values for each combination, the size of the differences, and uncertainty intervals.
Shortfalls occur mostly in OPF training domainsRelative performance appears to depend on the type of domain being evaluated.Which specific domains account for the exceptions and how much they affect each architecture.
04

Recall does not answer the whole question

In a redaction task, recall usually refers to the proportion of relevant sensitive information that the system successfully detects. It is an important dimension: failing to flag information that should have been protected can leave it exposed. But higher recall does not automatically show that redactions are correct in every case.

It also matters how much text is wrongly concealed. Over-redaction can remove useful information or make a document incomplete; a missed item can leave exposed content that should have been protected. The available summary does not provide figures for these two types of error or explain how they are balanced. It is therefore not possible to compare the practical cost of errors by ASIRF and OPF using the highlighted figure.

The evaluation should not be mistaken for a deployment test either. The inclusion of fictional domains does not demonstrate performance on real-world language, genuine personal data, differences between organizations, or high-risk situations. The preprint presents initial experimental evidence as described in a summary, not an independent validation for production use.

05

What to check before using it with real data

Before considering ASIRF for processing personal information, it would be necessary to examine complete and reproducible results: metrics broken down by domain and dataset, the behavior of both architectures, missed items and over-redactions, and the criteria used to construct out-of-distribution evaluations. The evaluation instructions and the way expert definitions were created and reviewed would also need to be understood.

The information provided does not confirm whether code, data, or instructions for reproducing the experiment are available. Their availability should not be assumed, and publication alone would not guarantee independent reproduction. This requires checking the materials associated with the work and confirming that they allow the conditions and metrics to be repeated.

For a real application, the system would also need to be tested on data and tasks representative of the intended context, access to inputs and outputs would need to be restricted, and a policy for handling uncertainty would need to be defined. These are evaluation and deployment questions that the preprint’s aggregate result does not resolve. The Spanish Data Protection Agency’s guidance on agentic AI provides general context and distinguishes agent learning from retraining a language model; it is not evidence about ASIRF’s results.

The most limited conclusion is that ASIRF explores a way to make detection and redaction depend on definitions retrieved for the relevant domain, without retraining for each one. The summary reports a recall advantage over OPF in most of the combinations considered, but does not provide enough detail to judge all error types, reproducibility, or suitability for processing real personal data.

Checklist for an independent evaluation

  1. 01Obtain complete metrics by model, architecture, dataset, and domain.
  2. 02Separate missed items from unnecessary redactions and review examples of each type.
  3. 03Verify whether the data, code, and instructions needed to repeat the experiment are available.
  4. 04Test the system in the intended context, with human-review criteria and a predefined response to uncertain cases.

Open questions

  • Complete results by dataset, model, domain, or architecture are not provided, nor is the size of each difference against OPF.
  • The paper’s definition and construction of out-of-distribution domains are not specified in detail.
  • Missed items and over-redactions are not reported separately.
  • The availability of code, data, or evaluation instructions for reproducing the results is not confirmed.
  • The summarized information does not describe the specific operations in each call of the multi-agent system or the exact redaction procedure.
06

Keep exploring

06

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction