Ilustración editorial para Alertas de ciberseguridad con IA: cómo resumir, correlacionar y escalar incidentes sin dejar que una predicción cierre el caso
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

The goal is not to automate final judgment

A security operations center receives heterogeneous signals: endpoint detections, sign-ins, privilege changes, email messages, network activity, application logs, and cloud events. Many are repeated, incomplete, or lack business context. In that setting, an AI system can help structure information, identify candidate relationships, and draft an initial summary. That does not demonstrate that an incident exists, establish its scope, or determine an appropriate response measure.

The distinction is operational. A model output may be a well-formed working hypothesis; a decision to close, escalate, or contain changes a case's state and can affect users, systems, and evidence. The workflow should therefore let the model accelerate preparation tasks, while risk-changing transitions remain governed by testable rules and the authority assigned to responsible people.

The central recommendation is to separate four jobs. First, normalize and enrich events without replacing their original content. Second, group signals that might belong to the same investigation. Third, create a summary in which every material claim can be traced back to the records that support it. Fourth, recommend or prepare an action without confusing that proposal with authorized execution.

This guide helps determine what to automate within a security operation. To select a platform or integration pattern, also review the Choose path; to assess design alternatives, use the Compare path; and to explore related concepts and controls, use the Discover path.

02

Task map: what AI can do and what it must not assume

Not every task carries the same risk. Extracting event fields, translating a message, suggesting labels, or summarizing a timeline are supporting operations. Deduplication and correlation introduce interpretation: if applied with excessive confidence, they can hide an independent signal. Prioritization affects the work queue. Closure and containment directly affect risk exposure and, in some cases, business continuity.

A prudent design expresses these differences through separate permissions. The AI component does not need permission to modify detection rules, close cases, isolate devices, block accounts, revoke sessions, change network configurations, or delete objects. If it participates in an action, it should do so through a structured request that passes through policy, receives approval when required, and uses a least-privilege connector.

NIST's profile for generative AI identifies risks related to ungrounded outputs and the need to govern, measure, and manage the use of these systems. In a SOC, that means plausible wording is not evidence. Observations must remain separate from inferences, and gaps must be visible rather than filled with confident language.

Recommended autonomy level by task

TaskPermitted outputRequired control
Normalize and extract fieldsStructured record with provenancePreserve the original event and source identifier
EnrichTags and external or internal contextMark source, lookup time, and missing results
Group signalsCandidate group and similarity reasonsDo not suppress alerts or close cases because of grouping
Summarize an investigationTimeline, facts, and uncertaintiesInternal links to records and analyst review
PrioritizeSuggested priorityRules for criticality, scope, and minimum evidence
Contain or closePrepared request, not an autonomous decisionExplicit policy, authorization, and execution record
03

Normalize without losing provenance or meaning

Correlation is more reliable when data shares a minimum structure. A common schema can receive events from different tools without removing the attributes needed to review them later. The OCSF project maintains an open schema with event classes, objects, and attributes that can serve as a normalization reference. Adopting a schema does not remove the need to retain the native record: vendor, rule, or agent details can be decisive during an investigation.

Each normalized event should retain at least an immutable internal identifier, source identifier and type, observed time and receipt time, the related asset and identity when available, the detector that generated it, the normalization result, and an internal pointer to the original content. It is also useful to state which values came directly from a source and which were derived through enrichment or calculation.

Time deserves specific treatment. Two signals are not necessarily simultaneous because their timestamps are similar: ingestion delays, different time zones, or unsynchronized clocks may exist. Rather than asking the model to resolve that ambiguity with a conclusive sentence, the workflow should display the available times and mark uncertainty when a reliable sequence cannot be established.

MITRE ATT&CK distributes data and representations for working with knowledge of adversary behavior. It can be used as an enrichment vocabulary or to support analytics, but assigning a technique to an event does not prove intent or confirm an intrusion. The label should remain a classification or hypothesis alongside the evidence that motivated it.

Alert intake process

  1. 01Receive the event and assign it an ingestion identifier.
  2. 02Preserve the original record and calculate or store an integrity reference according to internal policy.
  3. 03Map fields to the common schema without removing relevant native attributes.
  4. 04Label every field as observed, derived, or unavailable.
  5. 05Apply permitted enrichments and record the source, time, and result of every lookup.
  6. 06Provide the grouping engine with a provenance-bearing representation, not only summarized text.
04

Grouping signals proposes a relationship; it does not declare one incident

Deduplication seeks to prevent multiple notifications from the same detector about the same fact from overwhelming a queue. Correlation seeks to identify distinct signals that may be related. They are different problems. A repeated alert with the same detection identifier, asset, and time window may be a deduplication candidate. An anomalous sign-in, a privilege escalation, and data transfer from the same environment may justify a grouped investigation, but they are not duplicates.

Before merging cases, define observable criteria: matching identity or asset, documented temporal proximity, network or process relationship, a common detection rule, a campaign identified by the intelligence team, or a known technical dependency. The system must show which criteria were met and which are missing. Semantic similarity between two descriptions can help discover candidates, but it must not be the only basis for hiding an alert or lowering its priority.

Maintain a reversible relationship between the group and its members. This lets an analyst separate a signal after discovering that it belongs to another incident. It also prevents grouping from turning a high-criticality event into a secondary note within a large set. Grouped cases should retain their own statuses and owners where policy requires it.

05

Minimum evidence must be visible before prioritizing or closing

A useful case is not persuasive prose; it is a reviewable set of facts, inferences, and decisions. The interface may provide a short synthesis, but it must allow access to the original event and relevant queries or enrichments. Minimum evidence varies by alert type, although a shared baseline reduces omissions: source, event identifier, time, asset, identity, triggered rule, observations, known scope, actions already taken, and missing data.

Explicitly separate four fields. “Observed facts” includes only data attributable to a source. “Inferences” includes proposed relationships or explanations. “Missing evidence” lists what would prevent confirmation or rejection of the hypothesis. “Proposed action” explains a possible measure, its rationale, expected impact, and whether it can be reversed. This structure makes it harder for an inference to be presented as though it had been observed.

NIST's publication on security and privacy controls addresses audit-event generation, content, and review. For this workflow, the practical principle is that a decision must be reconstructable: which alert initiated it, what data was queried, which tool was involved, who approved it, and which operation occurred. Saving only the model's final summary is not enough.

Minimum contents of an alert case file

ElementQuestion it answersHandling
Original eventWhat happened according to the source?Retain an internal reference and preserved content
ProvenanceWho generated the data and when did it arrive?Record source, detector, and timestamps
Asset and identity contextWho or what may be affected?Distinguish observed attributes from enriched inventory data
System inferenceWhat relationship or explanation was proposed?Show confidence level and reasons
GapsWhat is not yet known?Do not replace them with a conclusion
Decision and approvalWho changed the case state?Record identity, time, and applied policy
Executed actionWhat actually changed?Store request, result, errors, and reversal if any
06

Decision tree for assisted triage

The tree should begin with workflow integrity, not with confidence expressed by the model. If the original event is missing, the alert lacks sufficient provenance, or the data is contradictory, the system may request more information or route the case to review; it must not conclude that risk is low. Absence of evidence is not evidence of absence, especially when telemetry is partial.

Some alerts must retain human review from the start. These include alerts involving critical assets, elevated privileges, possible lateral movement, exfiltration, impact on essential services, regulatory or legal obligations, and any case where a possible response could materially affect users or production. The specific list must derive from each organization's inventory, risk appetite, and obligations.

Cases with conflicting signals, insufficient confidence, possible propagation, or missing relevant observability must also be escalated. Prioritization may combine technical severity, business criticality, potential scope, and evidence quality, provided that the organization documents how those factors are calculated. A single score is useful for ordering work, not for concealing the components that produced it.

Decision path for each alert

  1. 01Is an original record accessible and is its provenance identifiable? If not, review or enrich; do not close automatically.
  2. 02Does the case affect an asset, identity, or service defined as critical? If yes, assign priority human review.
  3. 03Are there observable criteria for deduplication or grouping? If not, keep signals separate.
  4. 04Is the suggested priority supported by visible evidence, criticality, and scope? If not, mark it as provisional.
  5. 05Is the proposed action low-impact and reversible under policy? If not, require named human approval.
  6. 06Does the action change access, connectivity, data, or configuration? If yes, record authorization and result before updating the case.
07

Escalation and containment: a recommendation is not an order

Incident response should be connected to risk management, available resources, and recovery processes. NIST guidance on incident response provides a framework for considering response within that management, rather than as an isolated sequence of automations. In practice, the escalation policy should specify who receives a case, what information they need, and when a decision must be made.

Classify actions by impact and reversibility, not only by technical ease. Drafting a ticket or collecting evidence is usually low impact. Preparing a request to block an account or isolate a device does not yet alter the environment, but it should include the reason and possible consequences. Executing isolation, account blocking, credential revocation, firewall changes, deletion, or data quarantine can interrupt operations or complicate later analysis; it normally requires explicit authorization defined by policy.

Even an apparently reversible action needs conditions. Isolating an endpoint can interrupt a business process; revoking a session can stop a legitimate operation; blocking an indicator can have unexpected effects if it is shared. Policy must define who can approve, when an urgent exception is allowed, how it is documented, and what the rollback plan is.

Approval decision by impact

Action classExamplesProposed control rule
InformationalSummary, ticket, inventory queryMay be automated if traceability is retained
PreparationBlock draft, artifact collectionMay be automated without executing the change
Low-impact reversibleTag, case assignment, temporary increase in observationAllow only if a policy defines scope and reversal
High-impact or access-relatedIsolation, account block, revocation, firewallRequires named human authorization except under an approved emergency procedure
Destructive or hard to reverseDeletion, broad configuration changeRequires strengthened control and evidence of necessity
08

Treat emails, tickets, and external sources as untrusted data

An email message, ticket, alert description, or external page may contain instructions directed at an analyst or at the model. Those instructions may attempt to alter classification, request that evidence be ignored, induce a tool call, or request a response action. Retrieved content must be treated as potentially hostile evidence, not as an extension of SOC policy.

CISA's cybersecurity and AI collaboration playbook includes cases related to instruction injection. For an alert workflow, the defense does not rely only on telling the model to “ignore” those instructions. A technical boundary must exist: policies, permissions, and tool definitions come from trusted components; documents, emails, and tickets are labeled as untrusted content; and the model cannot turn encountered text into authorization.

Tool calls must be validated against schemas and allowlists of permitted operations. A read-only inventory or log query can be authorized differently from a modification at an identity provider. If external content asks for an operation, the system must record it as investigation data and apply normal policy; it must never execute it merely because it appears in text.

09

Test the workflow with real history and adversarial cases

Before increasing autonomy, evaluate the design with historical incidents and test sets separate from those used to tune rules or instructions. The goal is not to check whether the summary “sounds good,” but whether it preserves evidence, prioritizes usefully, and avoids harmful closures or groupings. When historical records contain sensitive data, apply the appropriate access and minimization controls.

Include common noise, incomplete telemetry, repeated alerts, asset-name changes, distributed campaigns, known false positives, and cases in which two similar signals proved to be independent incidents. Also add adversarial inputs with embedded instructions in emails, tickets, and log fields. For every case, define the expected result and prohibited behavior, such as closing without sufficient evidence or executing an action without authorization.

Tests must cover degradation. If asset inventory is unavailable, an enrichment fails, or the model cannot produce a valid output, the case should return to an explicit safe state: pending review, with the reason for missing data. A missing dependency should not be replaced with an estimate presented as fact.

10

Measure outcomes and retain a complete decision history

Metrics should reveal both efficiency and potential harm. Time to first useful investigation indicates whether the system helps an analyst get oriented. Prioritization accuracy shows whether important cases reach review earlier. Evidence coverage indicates how many material claims have accessible provenance. Correct escalations, false closures, reversals, and containment time reveal effects that an alert-volume metric does not capture.

Define “false closure” before measuring it. For example, it may be a closed case that is reopened or later linked to a confirmed incident within an agreed window. The window, confirmation criteria, and treatment of new telemetry must be documented; otherwise, teams may compare figures that mean different things. Also analyze results by source, asset type, and criticality level to identify where automation fails.

An investigation identifier should link the initial alert, grouped events, enrichment queries, model outputs, tool calls, approvals, and executed actions. The record must distinguish an action requested from an action completed. Any error, denial, or reversal must also be recorded. This traceability allows an incident to be reviewed without depending on a person's memory or on one generated summary.

Automation becomes mature when its limits can be examined. Periodically review samples of closed, escalated, and contained cases; compare the decision with the evidence available at that time, not only with what became known later. Adjust policies, rules, and tests when patterns of omission, over-grouping, or unintended action effects emerge. The objective is to reduce repetitive work while retaining the human ability to question a conclusion.

Post-incident review of an AI-assisted decision

  1. 01Retrieve the investigation identifier and all associated original events.
  2. 02Reconstruct the timeline of received data, enrichments, inferences, and state changes.
  3. 03Separate evidence available at the time of the decision from information learned later.
  4. 04Verify which policy permitted the closure, escalation, or action and who approved it.
  5. 05Assess whether the model output confused facts, hypotheses, or missing data.
  6. 06Record corrections to rules, permissions, regression tests, and review procedures.

Open questions

  • Specific priority thresholds, windows for detecting false closures, and the list of critical assets depend on each organization's risk, regulation, and architecture.
  • Whether an action is reversible depends on the environment: a technically reversible measure may have material operational consequences.
  • A normalization schema improves interoperability, but it does not guarantee that all sources provide the fields needed for investigation.
  • The provided sources support principles of risk management, response, traceability, normalization, and AI-risk handling; they do not establish a single autonomy policy applicable to every SOC.
11

Keep exploring

11

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction