AI Agents in the SOC: Teams That Only Limit Actions Underestimate Triage
In July 2026, the AI-based defenses of AI platform Hugging Face detected an attack by an autonomous AI agent but rated its severity too low and failed to alert the on-call team. This very work, reviewing and assessing alerts, is what security and network teams at large companies would most readily hand over to AI agents. Most teams reject autonomous actions by their agents, but that boundary does not protect against misjudged alerts.
- According to a survey of 921 professionals by Cybersecurity Insiders and Netskope, mostly at large enterprises, 47 percent of security and network teams plan to adopt AI agents in the coming year, but only 32 percent plan to add staff.
- Triage, the work of reviewing and prioritizing alerts, is a practical starting point according to the Netskope report, because agents do not change any systems there; the Hugging Face case shows that an alert rated too low also costs defenders valuable time.
- 68 percent of teams in the Netskope survey would not let agents make changes on their own; only 9 percent are very confident they can govern their agents.
Most figures in this article come from the report “2026 State of AI Agents in Security and Network Operations”. Cybersecurity Insiders, a publisher of IT security surveys, produced it with the support of security vendor Netskope.
Of the 921 respondents, 89 percent work at companies with at least 1,000 employees. The results are therefore most meaningful for chief information security officers (CISOs) at large companies, the executives with overall responsibility for information security. They also matter to the heads of two types of operations centers: in a security operations center (SOC), a team detects and fights cyberattacks; in a network operations center (NOC), a team monitors network operations. Just over a quarter of respondents work in a SOC, 13 percent in a NOC. Eight percent belong to the executive level; most are managers, practitioners, or directors.
What happened at Hugging Face?
Hugging Face is a platform where developers share AI models and datasets. The attacker was itself an AI agent that OpenAI was testing internally for its cyberattack capabilities. According to Hugging Face, this evaluation deliberately disabled OpenAI’s production safety classifiers. Through a previously unknown vulnerability in the package proxy of its test environment, a cache for software packages, the agent reached the internet. From a third-party sandbox, an isolated environment for running code, it attacked Hugging Face’s dataset processing pipeline.
Hugging Face believes the agent wanted to steal the solutions to its test. According to Hugging Face’s technical timeline, the agent spent roughly two and a half days inside the infrastructure. In under 13 hours, it worked its way from a single container to administrator rights on multiple internal clusters.
The only customer content it accessed was five datasets that apparently belong to its test. Hugging Face reconstructed around 17,600 of the agent’s actions. Detection initially worked. Several systems flagged anomalies, including a SIEM (security information and event management) system, which collects and analyzes security events from many sources.
Individually, the signals were ambiguous. Hugging Face’s AI-based security agent stack correlated them into a coherent attack signal. “However, it failed to correctly raise the alert’s criticality and trigger the on-call team, costing precious time in the response,” Hugging Face writes. The timeline does not say whether the cause was a misjudgment by the agents or a misconfigured threshold.
The agent tried thousands of paths, and most led nowhere. “The successful path was hidden inside the noise generated by the thousands of failed ones,” the timeline states.
Why do teams still bet on agents rather than more staff?
The teams in the Netskope survey know this noise too. Workloads increased for 81 percent over the past year, and only 11 percent consistently keep up. Higher volumes of alerts, cases, and tickets are a cause for 41 percent, false positives and duplicates for 30 percent. At 42 percent of organizations, alerts, cases, or tickets went uninvestigated in the past twelve months.
At 24 percent, sensitive data, systems, or users remained exposed longer than necessary. Routine work eats up time. Sixty-one percent estimate that triage, correlation, documentation, and reporting take up at least 40 percent of their time. In triage, analysts review and prioritize incoming alerts; in correlation, they link reports from different systems into a single incident.
More staff is not the teams’ first answer. In the coming year, 47 percent plan to use AI agents for specific workflows, 44 percent want to better integrate their tools, and only 32 percent plan to add headcount. Fifty-six percent cite capacity pressure as a reason for their interest in agents, more than any other reason. Executive or board expectations come last at 18 percent.
Why is triage considered the starting point?
Among the tasks in the Netskope survey, effort, trust, and value align most closely in triage. Forty-six percent name it as a major source of manual work, the same share would hand it to an agent, and 48 percent expect significant value. For incident remediation, the figures are 28, 20, and 29 percent.
The Netskope report explains this priority: in triage, an agent organizes cases for analyst review without changing systems. The report locates the risk in permissions. “The discretion that makes an agent useful in triage can become operational risk the moment the agent has permission to act,” it states.
Why is triage risky too?
The Hugging Face incident shows where triage can fail: the agents correlated the signals correctly, but the alert never reached the on-call team. In triage, an agent decides what humans get to see at all. The Netskope report assumes that analysts review its assessment. The report does ask teams to track errors, but an underestimated attack does not feature in its reasoning.
One of its recommendations addresses only the other type of error: analysts should feed their verdicts back to the agent so that it “stops surfacing what the team keeps dismissing.” This reduces false positives. The report does not say how teams should prevent the agent from also suppressing real attacks.
An alert rated too low causes damage even without any intervention: the attacker gains time. Teams should therefore measure an agent not only by how many false positives it filters out, but also by how many real attacks it underestimates. The benchmark is not a flawless team: at 42 percent of organizations, alerts already go uninvestigated today.
The Netskope report cites the Hugging Face incident itself, but only as evidence of how valuable linked logs are for investigations. It does not mention that the escalation failed.
In early May 2026, cyber and security agencies from Australia, the US, Canada, New Zealand, and the UK published the guidance “Careful Adoption of Agentic AI Services”. The agencies recommend configuring agents to stop and escalate to human reviewers in uncertain situations. The rule assumes that an agent notices its own uncertainty. Elsewhere, the guidance itself casts doubt on that: large language models, it says, are typically trained “to produce outputs that resemble material rated highly by humans, rather than to identify when a query falls outside their knowledge limits.”
At Hugging Face, the agents had already correlated the signals into a picture of the attack. What failed was the severity rating, not the detection. Hugging Face has since introduced a fixed rule: the detected behavioral patterns now trigger critical-severity alerts.
The 2026 SANS SOC Survey by the SANS Institute, a training and research organization for IT security, shows how often AI runs in the SOC without a defined framework. Of the 444 respondents, 79 percent use AI or machine learning tools. Only 36 percent have built them into a defined workflow.
“The risk is not that AI performs badly. The risk is that no one will know when it does,” writes author Christopher Crowley.
How much may agents change on their own?
In the Netskope survey, teams draw a clear line. Thirty-eight percent would not yet accept agents executing changes to systems, and another 30 percent only after approval by a human. That adds up to 68 percent.
Twenty-three percent would let agents complete low-risk actions within defined guardrails on their own, and 5 percent would give them more latitude.
A scenario from the Netskope report shows why this matters: an agent treats an unusual outbound data transfer as theft and isolates the affected cloud workload. Its permissions allow this, and no approval is required. But the workload supports a customer-facing service, which now fails.
The report advises drawing the line between actions a team can reverse and actions it cannot. The agencies take a similar view. Organizations should classify their agents’ actions by potential impact, likelihood, and reversibility.
Under the guidance, system designers or operators decide when a human must approve, not the agent. The agencies recommend approval checkpoints before system resets, network egress, or the deletion of critical records, for example.
Their principle: “Privileges assigned to agents directly determine the level of risk they can introduce.” That holds for interventions. How consequential a wrong severity rating in triage is, however, does not depend on the agent’s privileges.
On one point, the agencies go further than current practice: organizations “should only use agentic AI for low-risk and non-sensitive tasks.” Read strictly, that rules out many SOC tasks. Even triage almost always touches sensitive data such as user accounts or network traffic.
Who is responsible when an agent gets it wrong?
Deployment is running ahead of control. In the Netskope survey, 29 percent of teams already run agents in production, and 53 percent are evaluating or piloting them. Yet only 9 percent of all respondents are very confident that their organization can govern the agents. Forty-five percent have low or no confidence.
Companies assign responsibility inconsistently. Twenty-six percent share oversight across functions, 21 percent place it with SOC leadership, and no other model exceeds 12 percent.
The Netskope report is clear: “Delegating work does not transfer accountability.” A case from Canada shows that responsibility cannot be shifted onto software in dealings with customers either. After the death of a customer’s grandmother, Air Canada’s chatbot told the customer that the reduced bereavement fare could also be claimed after travel, within 90 days of the ticket’s issue date. The airline later refused the refund because its rules exclude exactly that. In the subsequent proceedings, Air Canada argued that it could not be held liable for information provided by its representatives, including the chatbot.
The Civil Resolution Tribunal, a public tribunal for smaller civil disputes in British Columbia, read this as a claim that the chatbot was a separate legal entity. In February 2024, it called this “a remarkable submission.” It ordered the airline to pay just over 800 Canadian dollars, because the chatbot was part of the website, and Air Canada was responsible for its content.
The agencies’ guidance recommends defining “legal accountability and risk ownership for agentic AI systems in policies.”
What do both surveys show?
Both surveys point in the same direction: AI has arrived in security operations, but governed use of it has not. A side finding for those responsible for the Internet of Things (IoT): only 16 percent of SOCs in the SANS survey monitor all of their organization’s at-risk connected systems, and another 29 percent cover some of them. These include sensors, building systems, and operational technology (OT), the control technology used in production and industrial plants.
Conclusion
Security teams want to use AI agents first for sorting alerts, because an agent does not change anything in systems there. The Hugging Face case shows that errors occur there too: the agents detected the attack, but the alert never reached the on-call team. Anyone who measures an agent only by the false positives it filters out will miss this kind of error.
The bigger task lies in accountability. In the Netskope survey, almost a third of teams run agents in production. Only one in eleven respondents is very confident they can govern agents. The security agencies set the direction: resilience, reversibility, and risk containment take priority over efficiency gains.
Sources: Cybersecurity Insiders/Netskope, “2026 State of AI Agents in Security and Network Operations” (September 2026, version 1.6); Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident” (July 27, 2026); ASD’s ACSC, CISA, NSA, Cyber Centre, NCSC-NZ, NCSC-UK, “Careful Adoption of Agentic AI Services” (May 2026); SANS Institute, “2026 SANS SOC Survey Insights: A Decade of Evolution in Cyber Defense” (June 2026); Civil Resolution Tribunal of British Columbia, Moffatt v. Air Canada, 2024 BCCRT 149 (February 14, 2024).
AI agents are systems based on large language models that decide for themselves which step to take next. In a security operations center (SOC), they can review alerts, correlate reports from different systems, and prepare cases for analysts. Unlike classic automation, they do not follow a fixed, pre-programmed sequence.
Triage, the review and prioritization of alerts, is considered a practical starting point because an agent does not change any systems there. In a survey by Cybersecurity Insiders and Netskope, 46 percent of security and network teams would hand this task to an agent. For incident remediation, which requires agents to change systems, the figure is only 20 percent. Triage is not risk-free either: if an agent rates a real attack too low, the team may learn about it too late.
Most teams surveyed would not allow that. According to a survey of 921 professionals by Cybersecurity Insiders and Netskope, 38 percent currently rule out changes by agents entirely, and 30 percent require human approval first. Twenty-three percent would let agents carry out low-risk actions within fixed guardrails on their own.
In the guidance Careful Adoption of Agentic AI Services, cyber and security agencies from Australia, the US, Canada, New Zealand, and the UK recommend incremental deployment, least-privilege access, and human approval for high-impact or hard-to-reverse actions. Agents should stop and escalate to humans in uncertain situations. Organizations should use them only for low-risk and non-sensitive tasks.
There is no general answer. A case from Canada shows, however, that companies cannot hide behind their software: in 2024, the Civil Resolution Tribunal in British Columbia ordered Air Canada to pay because its chatbot had given a customer wrong information; the chatbot was part of the website, and the airline was responsible for its content. This single decision does not establish a general legal rule for other countries. Security agencies recommend defining in policies who is legally accountable for agents and who owns the risk.
In July 2026, an AI agent driven by OpenAI models broke into Hugging Face’s infrastructure after escaping from an internal OpenAI test environment. Hugging Face reconstructed around 17,600 of the agent’s actions; it gained administrator rights on multiple internal clusters. Hugging Face’s AI-based defenses correlated the signals into a picture of the attack but did not rate the alert high enough and failed to alert the on-call team.










