Ilustración editorial para Incidentes graves de IA: cómo preservar evidencia, contener el daño y decidir si un fallo debe notificarse bajo el artículo 73 del AI Act
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

A failure is not automatically a serious incident

Responding to a dangerous outcome from an AI system begins by separating concepts that are often conflated. A model error may be an inaccurate, inconsistent, or unwanted output. An operational incident may involve a service outage, incorrect configuration, failed integration, or tool use outside intended behaviour. Harm is a specific negative consequence for a person, organisation, property, the environment, or a process. A serious incident, in the regulatory sense relevant to this guide, is a narrower category linked to specific effects defined by the AI Act.

Not every hallucination, quality reduction, user complaint, or offensive response independently triggers the Article 73 regime. Nor is it prudent to dismiss a case because the isolated output appears minor: an apparently ordinary response may have influenced a clinical, recruitment, access-to-service, physical-safety, or infrastructure-operation decision. The assessment must start from observable facts, rather than internal labels such as “minor bug” or “non-critical complaint.”

The consolidated AI Act text defines a serious incident through outcome categories including death or serious harm to health, serious and sustained disruption of the management or operation of critical infrastructure, infringement of obligations intended to protect fundamental rights, and serious harm to property or the environment. Classification requires examination of the event, the affected system, and the possible link between them. It should not be replaced by an internal severity score that has no documented correspondence to those categories.

It is useful to maintain two tracks from the first alert. The technical track seeks to stop unsafe behaviour and restore a controlled service. The evidence-and-compliance track seeks to preserve facts, assess the possible causal link, coordinate provider and deployer, and determine whether a communication to the competent authority should be prepared. They can proceed in parallel, but a rushed correction can impair the second track if it changes or removes information needed to explain what occurred.

Initial decision tree: four questions before labelling the case

QuestionIf yesIf no
Could the system be subject to the applicable high-risk regime?Open a legal and technical scope assessment; identify the classification route and responsible operator.Do not assume Article 73 applies; preserve evidence and review other applicable contractual, sectoral, or safety duties.
Is there a verifiable event involving harm, materialised risk, or a dangerous outcome?Open the incident record and activate proportionate containment.Record the signal as an anomaly, with a reassessment threshold if new evidence emerges.
Does the event fall, or could it fall, within a serious-incident category?Escalate to compliance and legal teams; assess causation or its reasonable probability.Do not present it as a serious incident; document the reasons and continue the technical investigation.
Is a link with the system known or reasonably probable?Prepare the reporting clock and preserve the relevant state before further changes.Keep hypotheses open; do not assert causation without sufficient evidence.
02

Scope: identify the system, the operator, and the high-risk route

Article 73 addresses providers of high-risk AI systems that have been placed on the Union market. The first fact to verify is therefore not the severity perceived by the team, but the exact system involved, who is its provider for the purposes of the Regulation, and whether it was subject to the relevant high-risk classification. Article 6 of the AI Act provides two main classification routes: certain systems related to regulated products or safety components, and the use cases listed in Annex III, subject to the conditions and exceptions established in the Regulation itself.

It is not enough to state that a foundation model, conversational interface, or automation is “high-risk” because of its subject matter. The unit of analysis must be the system made available or used in the specific context: its version, intended purpose, integration, enabled functions, users, input data, and the outcome it produced. The same component can form part of configurations with different obligations. The Commission’s draft guidance on classification can help structure the review, but it remains a draft and does not replace the binding text or case-specific legal assessment.

Planning must distinguish preparation from actual applicability. Organisations can build their logging, preservation, contact, and escalation processes now. However, the application dates for particular high-risk categories must be checked against the legally applicable version of the Regulation and any amendments in force when deciding a case. An internal planning date should not be turned into a conclusion that an obligation is already enforceable.

To navigate related work, a team can distinguish this protocol from general Safety reviews, assessment of alternatives in Compare, and capability discovery in Discover. Those activities may provide context, but an incident investigation requires a record focused on the event that occurred and the configuration that was actually deployed.

03

The first hours: preserve before correcting

When a potentially relevant case is detected, appoint an incident lead with authority to coordinate operations, product, security, quality, and compliance. Their first operational duty is to establish a timeline: when the event occurred, when it was detected, who received each alert, which systems remained active, and what decisions were made. The time at which the provider became aware and, where applicable, the deployer became aware should be recorded separately, together with the evidence supporting each time.

Preserving evidence does not mean indiscriminately copying all available data. It means retaining, proportionately, with integrity and controlled access, the elements that permit reconstruction of the investigated behaviour. At a minimum, the record should link request and session identifiers; inputs and outputs; model version and parameters; system prompt and templates; policies; retrieval configuration; retrieved documents or their identifiers; tool calls and responses; human approvals; the identity or role of the operator; and configuration changes near the event.

Preservation should include integrity metadata: source, extraction date, responsible person, export method, file hash where feasible, access controls, and every subsequent transformation. Where personal data, trade secrets, or security information are involved, access should be restricted under the applicable rules. Restricting access does not justify deleting elements needed for investigation. If a datum cannot be retained, the record should explain what was removed, why, when, and what alternative evidence remains.

The AI Act requires the provider to investigate the serious incident and the related system immediately, including risk assessment and corrective measures. It also provides that the system should not be altered before reporting where the alteration could affect the subsequent assessment of the incident’s causes. In practice, containment must therefore be designed to limit harm without erasing the state that needs to be examined.

First four-hour process

  1. 01Open a unique incident identifier and record the trigger, detection time, and alert source.
  2. 02Appoint the lead, deputy, and decision channels; keep the factual log separate from hypotheses and assessments.
  3. 03Protect logs, configurations, deployment artefacts, and evidence of external actions through controlled copying and a chain-of-custody record.
  4. 04Apply a reversible risk-reduction measure where possible, such as disabling a tool, blocking a specific flow, or requiring human review.
  5. 05Record every containment change, its owner, affected scope, and verification that it did not destroy evidence.
  6. 06Escalate to compliance and legal teams if the case may concern a high-risk system or a serious-incident category.
04

Contain harm without disabling the investigation

Containment has no single correct form. Suspending the entire system may be necessary where a serious risk persists, but it may also affect essential processes or push users toward uncontrolled alternatives. Other measures may be more proportionate: temporarily removing an execution tool, reducing permissions, preventing high-impact automated actions, blocking an identified set of inputs, disabling a compromised retrieval source, or requiring a human check for a specific decision.

Every measure should address an explicit harm hypothesis and have review conditions. “Put the system into safe mode” is not a verifiable description unless it defines what capability was blocked, what remained available, which population was affected, what alternatives were offered, and how the effect was checked. Communication to customers, users, and operators may form part of containment, but it should be based on confirmed facts and should not attribute causation before the investigation supports it.

Rolling back to an earlier version requires caution. It may resolve an immediate symptom, but it does not prove that the cause lay in the withdrawn version and may introduce differences that make the case impossible to reproduce. Before changing anything, preserve the deployed artefact, configuration, and records needed to compare the prior and subsequent states. If change is unavoidable to prevent harm, document the necessity and scope of the deviation.

Market-surveillance authorities may have their own powers regarding products presenting a serious risk, including assessment measures and restrictions. Those actions do not replace the internal investigation or turn the organisation into an authority. The team should be ready to provide verifiable facts, assessments, and applied measures if it receives a request.

Proportionate containment matrix

Observed situationPossible measureEvidence to preserveReview criterion
A tool may carry out an incorrect external actionDisable the tool or reduce permissions to read-onlyRequest, parameters, response, authorisation, and external action or attempted actionNo requests remain pending, the scope has been identified, and a controlled test of the correction exists
The output may influence a high-impact decisionRequire human review and block the affected automated decisionOutput, information shown to the reviewer, identity or role, final decision, and rationaleRisk assessment validates that the restored flow retains effective controls
Retrieval supplies erroneous or unauthorised documentsIsolate the affected index, corpus, or connectorDocument identifiers, index version, queries, and rankingSource, permission, and retrieval behaviour have been reviewed
There is immediate, unbounded riskSuspend the affected capabilityDeployment status, affected population, alerts, and urgency rationaleA competent internal authority approves resumption based on evidence
05

Reconstruct the case as a chain of decisions

A useful investigation must answer a simple question: what exact system produced or contributed to the outcome, and through what sequence? A conversation transcript is not enough. A reproducible record links the case identifier to the model and its version, inference parameters, system instructions, assembled prompt, attached data, retrieval configuration, selected documents, available tools, calls made, active policies, and human decisions.

Facts should also be separated from the assumed mechanism. Fact: a tool received specific parameters and a recorded action occurred. Hypothesis: the model misinterpreted an ambiguous instruction. Fact: an operator approved a recommendation. Hypothesis: the interface did not display enough context to detect the error. This separation prevents the first diagnosis from becoming the final narrative and permits review of alternative hypotheses, including data, integration, interface, training, human-process, or external failures.

Complete reproducibility will not always be possible. An external service may have changed, an input may be missing, an outcome may depend on randomness, or data retention may be limited. In those cases, the record should precisely describe the gap, its effect on conclusions, and the substitute tests used. Failure to reproduce does not by itself demonstrate that the system was unrelated to the incident.

Reconstruction must include changes made during the response. Without this record, a later improvement can be confused with the original configuration, and a safe outcome obtained after containment can wrongly be presented as evidence that the incident was impossible. Maintaining comparability among the initial state, contained state, and corrected state is a practical condition for learning from the case.

06

Classify severity and causal link without delivering a verdict too early

Classification should begin with specific questions. Was there death or serious harm to health? Did the management or operation of critical infrastructure suffer serious and sustained disruption? Was there an infringement of obligations intended to protect fundamental rights? Was there serious damage to property or the environment? For each question, the record should identify the alleged event, supporting sources, known extent, affected people or assets, and remaining uncertainties.

The link with the system must then be assessed. Article 73 does not require an organisation to wait for definitive proof of causation where there is a reasonable probability of a link, but it also does not allow mere temporal coincidence to substitute for analysis. The file should document plausible causal pathways, supporting evidence, alternative factors, and pending tests. For example, an incorrect output may be relevant, but the final decision may have depended on independent human review or inaccurate external data; both elements need investigation.

An internal severity matrix may facilitate escalation, but it should not replace the legal definition. The internal form should contain separate fields for “observed impact,” “risk of additional impact,” “potential regulatory category,” “confirmed link,” “reasonably probable link,” and “link not established.” This makes clear which part is factual and which is provisional analysis.

A decision that a case does not meet the threshold should be reasoned and reviewable. New information may come from a deployer, user, connected tool, or authority. Closing the reporting assessment does not mean closing the technical record or deleting the evidence.

07

Provider and deployer: coordinate information rather than transferring the problem

Provider and deployer may hold different parts of the evidence. The provider will commonly control system documentation, versions, testing, operational logs, and product corrective measures. The deployer may know the use context, affected population, human decisions, local data, material consequences, and received communications. A contract may allocate operational tasks, but it should not prevent rapid delivery of information needed to meet applicable duties.

Prepare a contact matrix and incident operating clause before a case occurs. It should include out-of-hours contacts, minimum information categories, secure channels, internal deadlines shorter than regulatory maxima, preservation rules, a procedure for approving communications, and treatment of protected data. The purpose is not to transfer responsibility automatically, but to reduce the time between awareness of an event and obtaining the elements needed to assess it.

The deployer should preserve and facilitate data under its control that are necessary, within applicable legal limits. The provider should not demand perfect reproduction as a condition for starting assessment. Conversely, the deployer should not make local changes, erase logs, or communicate a technical cause as confirmed without coordinating the record. Where several entities are involved, a shared evidence-request log helps distinguish what has been delivered, what is pending, and what cannot be obtained.

Transparency, logging, and cooperation obligations associated with high-risk systems may be relevant to making this coordination work, but they do not automatically make the deployer responsible for the provider notification contemplated by Article 73. Final allocation depends on the effective role of each entity and the circumstances of the investigated system.

Minimum information exchange between provider and deployer

PartyMain contribution to the recordRisk to avoid
ProviderVersion identification, technical documentation, available logs, system analysis, risk assessment, and corrective measuresWaiting for all external data before preserving and analysing its own evidence
DeployerUse context, affected users, human decisions, local records, observed consequences, and local measuresChanging the flow or deleting data before reporting the change
BothShared timeline, evidence requests, containment status, hypotheses, and communication of changesPresenting incompatible conclusions or withholding relevant information because no agreed channel exists
08

Notification: manage the clocks without turning them into an automatic formula

Article 73 establishes a duty to communicate serious incidents to the market-surveillance authorities of the Member States in which the incident occurred, under the conditions set out in the Regulation. Communication is connected both to awareness of the incident and to determining a causal link or its reasonable probability. The team should therefore record separately the time of awareness, the time at which a provisional conclusion on the link was reached, and the evidence supporting both milestones.

The Regulation provides maximum periods of two, ten, and fifteen days for different situations, as well as the possibility of an incomplete initial report followed by additional information. The exact allocation of each period depends on the specific incident category and the wording applicable to the case. It should not be inferred from an abbreviated internal table or technical severity alone. The protocol should trigger immediate legal review, calculate the deadline from a documented milestone, and record why that clock was chosen.

An initial report should not be filled with artificial certainty. If the cause is unknown, state that it is under investigation, describe the confirmed facts, known scope, containment measures, available evidence, and plan for completing the information. The possibility of submitting an incomplete report does not justify delaying a required communication or omitting a diligent investigation.

The Commission’s published template and guidance on serious incidents are useful reference materials for organising fields and sequence, but they are presented as drafts subject to consultation. They should not be treated as final binding guidance. In a real case, the team should check the communication content against the applicable Regulation and the instruction of the competent authority.

Operational control of the notification clock

  1. 01Record the initial event and the date and time at which each entity became aware of it.
  2. 02Verify high-risk status and the provider role without delaying preservation and containment.
  3. 03Provisionally classify the outcome against serious-incident categories and document uncertainties.
  4. 04Assess and record causation or the reasonable probability of a link, including alternative hypotheses.
  5. 05Request legal review to assign the two-, ten-, or fifteen-day period and determine recipient authorities.
  6. 06Where appropriate, prepare a factual initial report and a plan with owners and dates for completing it.
  7. 07Record every communication sent, acknowledgement received, later update, and related corrective measure.
09

Investigation, correction, and controlled resumption

The investigation does not end when a communication is sent. It should explain the cause or contributing causes with confidence proportionate to the evidence: model behaviour, input data, retrieval, tool, interface, permissions, configuration, human oversight, training, operational process, or a combination. A single root cause can be misleading where the incident depends on several controls that failed or were absent.

Corrective action should be linked to the identified harm pathway. Adjusting a prompt may be insufficient where the problem was excessive tool permission; adding human review may be insufficient where the reviewer lacks required data; removing one document may be insufficient where the connector continues to ingest unauthorised sources. For every measure, define the risk it reduces, residual risk left behind, possible side effects, and how it will be validated.

Validation should include the incident case and broader regression testing. Testing only the original conversation can encourage an overly tailored correction. Prefer a combination of representative tests, edge cases, permission tests, and end-to-end flows through the external action, together with review of effects on users and affected groups where appropriate. Retain results, test environment, versions, and acceptance criteria.

Resumption should not be an implicit decision made when a ticket is closed. It should have an owner, authorisation criteria, initial scope, enhanced monitoring metrics, a rollback mechanism, and conditions that require suspension again. If material uncertainty remains, the organisation may keep the affected capability limited while completing the investigation. That decision and its rationale should be recorded.

10

Quarterly preparation: verify that the protocol works before an incident

Response capability is not demonstrated merely because technical logs or a written policy exist. The organisation must test whether it can recover a realistic case without relying on one person or tools that do not retain the necessary data. A quarterly exercise can select a risk-relevant flow, simulate an alert, and measure how long it takes to identify the deployed version, isolate the affected capability, obtain deployer records, and produce a timeline with verifiable sources.

The review should include changes in providers, models, tools, corpora, permissions, owners, and deployment markets. A protocol prepared for a static model may fail with model routing, phased deployments, dynamic retrieval, or third-party tools. It is also useful to check whether agreements with customers and suppliers enable rapid sharing of needed evidence with suitable confidentiality and data-protection controls.

Every exercise should result in observable improvements: missing log fields, decisions with no owner, inability to retrieve configurations, no emergency channel, or ambiguous suspension criteria. The aim is not to declare general compliance with the AI Act, but to reduce uncertainty in a concrete incident response. The most useful preparation makes it possible to say what is known, what is not known, and what was done to prevent harm from continuing.

Quarterly readiness checklist

  1. 01Check that every potentially relevant system has an owner, identified provider, intended purpose, and escalation contact.
  2. 02Run a log recovery for a test request and verify model version, prompt, tools, retrieval, and human decision.
  3. 03Test a reversible containment measure and document its operational impact and reversal.
  4. 04Review access to the evidence repository, retention, integrity, and custody procedure.
  5. 05Update the provider-deployer matrix, contacts, and secure communication channels.
  6. 06Submit a simulated alert to compliance and legal review to validate classification, causal link, and deadline control.
  7. 07Record deficiencies, assign owners, and verify closure in the next exercise.

Open questions

  • High-risk classification depends on intended purpose, use context, and the legally applicable text; the Commission’s draft guidance is not binding.
  • This guide does not independently assign every two-, ten-, or fifteen-day period to a factual category. That determination requires checking the case against the Article 73 in force and obtaining legal review.
  • The existence of causation or a reasonable probability of causation depends on the evidence of the case; it cannot be inferred from temporal proximity alone.
  • Application dates for obligations affecting specific system categories should be verified in the applicable version of the Regulation and amendments in force at the time of the incident.
  • Authority measures, sectoral duties, data protection, confidentiality rules, and national obligations may add requirements not developed in this guide.
11

Keep exploring

11

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction