AI SOC analyst: what to demand before you trust one
Hand an AI SOC analyst your tier 1 triage only if it shows what it ruled out, cites evidence you can check, drops its confidence score in the open, and keeps a human in the loop on every destructive action.
An AI SOC analyst is software that reads the alerts and cases your detection stack already produces and does the first pass a tier 1 analyst would: what it thinks happened, and why. Whether one deserves that work has almost nothing to do with which model sits behind it.
It comes down to six properties you can test in an afternoon — whether it shows the innocent explanations it ruled out, whether every value it cites is checkable against your own alerts, what it does when the evidence does not support its verdict, where your data goes, whether you can refuse a cloud model outright, and who has to approve anything destructive.
Most evaluations never reach any of that, because most demos are built to show a verdict. A verdict is the least interesting thing an AI analyst produces. Your team already has opinions; what the on-call analyst lacks at three in the morning is a defensible reason to act on one.
What an AI SOC analyst is for
It does not detect. Detection is what you already bought: a SIEM or XDR that fires alerts, a ticket queue that holds the backlog, and a person joining them together by hand.
That last step is the expensive one. CardinalOps read live production SIEM configurations for its 2025 State of SIEM Detection Risk and found detections covering 21% of MITRE ATT&CK techniques where the data already ingested could cover 90%, with 13% of the rules that did exist silently broken. Microsoft and Omdia, in their vendor-commissioned State of the SOC 2026, report that two thirds of SOC teams lose about a fifth of weekly capacity to manual correlation.
Two problems sit inside that, and only one is an AI analyst's. Missing and broken rules are detection engineering, and no model writes those for you. The manual correlation is the half software can take.
So the work an AI SOC analyst should take away is the tier 1 half: reading the alerts on a case, proposing what connects them, qualifying it, and saying what would make it wrong. Escalation and the decision to act stay with your analysts.
A tool that scores every alert on its own, without the case around it, is doing the cheap part of triage. That grouping should be deterministic and explainable before any model sees it, which is what a correlation engine is for.
Six things to demand from an AI SOC analyst
| Demand | The failure it prevents |
|---|---|
| The benign explanations it ruled out | A confident verdict that never considered the boring answer |
| Every cited value traceable to an alert on the case | Fabrication that reads as precision |
| A confidence score that comes down in the open when the evidence is thin | Certainty that survives its own missing facts |
| Permission to answer "not enough evidence" | A tool that must answer, so it always does |
| A local model option and one hard egress switch | Your incident data leaving with nobody's approval |
| An approval gate on every destructive action, and an execution trace | An automated mistake at estate scale |
Take them in order. The first three are about the quality of the reasoning; the last three are the guardrails.
Show me the innocent explanations you ruled out
A verdict arrived at without alternatives is a guess with formatting. Ask for the alternatives explicitly: which benign explanations were on the table, and what evidence removed each one.
Order matters more than most buyers realise. An analyst who states a conclusion and then assembles reasons is rationalising, and a model that writes its verdict first does the same thing — everything after that sentence is shaped to support an answer already chosen. Demand that the hypotheses and the evidence on both sides come before the verdict, so your analyst watches the argument being built rather than decorated.
Then test it on a boring case — the sort that makes up most of what a shift closes as a false positive. Bring a scheduled vulnerability scan, a backup job, an administrator doing administrative things.
A good AI SOC analyst names the innocent explanation first and tells you what would have made it the wrong one. A weak one turns a maintenance window into an intrusion narrative, because that is the text it has read most of.
Every value it cites must be checkable against your own alerts
The dangerous failure is not a vague answer. It is a precise one: a hostname, an address, a file path, all plausible, none of them present in the case.
Demand that the check is mechanical rather than the analyst's job. Every specific in the argument should resolve to an alert actually attached to the case, and a claim that resolves to nothing should be marked where the analyst can see it.
If the vendor's answer is that analysts will notice, they will not. Alert fatigue is exactly the condition in which a plausible hostname goes unchecked, and noticing is the work you are trying to give back.
This is also why the source matters less than the case. Whether the alerts came from Splunk, Microsoft Sentinel, CrowdStrike, Microsoft Defender, Cortex XDR, Elastic, OpenSearch, ClickHouse or Wazuh, the evidence check is the same operation: does the claim appear in what the case actually holds?
An AI analyst that only reads one vendor's alerts is not a triage layer. It is that vendor's lock-in with a language model on top.
Ask what happens when the evidence does not support the verdict
Every AI SOC analyst will reach a case where the argument does not hold up. The question is what it does then, and there are only three real answers.
- It stays confident. The worst outcome, and the most common, because confidence is a property of the writing style rather than of the evidence.
- It hedges quietly. Slightly better, and still useless. A hedge buried in the third paragraph is not a signal.
- It downgrades in the open. The verdict and its confidence score come down in front of the analyst, with the reason stated.
Demand the third, and demand that abstention is allowed. A tool obliged to produce a verdict on every case will produce one on every case, including the ones where the honest answer is that the evidence is not there. In our own product this is what SIROC does: an unsupported verdict is downgraded visibly, and the analyst is told why.
Where the data goes, and the right to refuse a cloud model outright
Ask three questions, and listen for properties rather than promises.
Can it run entirely on your own infrastructure? Not a private endpoint at somebody else's provider, but a model on your own hardware, on ordinary processors, with no key to buy. If the only architecture on offer sends your incident data out, then "we do not train on your data" is the strongest assurance you will ever get, and it is a promise rather than a property.
Can you refuse cloud calls outright? There should be one switch, set once for the whole deployment, that nothing underneath it can loosen. A per-user preference is not a control.
If you do allow a cloud model, what leaves? Demand redaction of secrets and personal data before the call, and a tamper-evident record of every outbound call — what left, where it went, what was stripped — that your data protection officer can export. Recollection is not evidence.
Then ask for two limits: a spend cap, and a breaker that stops the model until a person restarts it.
Our own commitments are on the security and trust page, including what we do not hold: no third-party pentest report, no certification, no SLA.
A person approves anything destructive
Isolating a host, resetting credentials, blocking an address at the perimeter: these are the actions that turn a wrong verdict into an outage. Demand an approval gate on every one, and then demand something harder — that no configuration can remove it.
Demand an execution trace with it: what ran, on whose approval, and what it changed. Human-in-the-loop is not a state of mind. It is a record somebody can read the morning after.
Vendors call this bounded autonomy. The phrase covers two very different things — a setting somebody switched on for you, and a ceiling nobody can raise. Ask which one you are being sold, and ask what enforces it.
If the answer is a checkbox on an administration screen, the control is only as strong as the least careful administrator you will employ over the next three years.
Gartner is right: there will never be an autonomous SOC
Gartner's formal position, published under exactly that title, is that there will never be an autonomous SOC. We agree with it, and we say so in our own material.
And the agentic SOC vendors publish the same refusal. CrowdStrike's Charlotte AI FAQ asks whether it replaces human SOC analysts and answers "No"; Palo Alto's agentic page carries a section headed "Maintain Absolute Control Over Agent Autonomy".
So treat agentic as a description of the plumbing rather than a promise about headcount, and treat the refusal as a buying criterion rather than a caveat.
A vendor selling you an analyst-free SOC is selling you an accountability problem: under NIS2 and DORA the clocks run against decisions your organisation made, and "the model classified it" is not a defence anyone will accept.
The useful goal is not fewer people. It is fewer cases that need one, and a written argument attached to each of the cases that do — which is also what survives a shift handover, where a verbal opinion does not.
How to run the evaluation in an afternoon
Bring three of your own cases: one true positive you have already resolved, one boring false positive, and one that was genuinely ambiguous at the time.
- Ask for the reasoning before the verdict on all three, and check whether the boring one is called boring.
- Remove a piece of evidence from the ambiguous case and run it again. The verdict should move, and the tool should say what moved it.
- Cut outbound internet access and run all three once more. Whatever still works is what you are actually buying.
- Ask, in front of the room, what it would take to make a destructive action automatic. Watch how long the answer takes.
We are pre-GA, with no customers to point you at and no accuracy figure of our own, because no cross-vendor benchmark exists that would make one meaningful. What we will do is put SIROC in front of your own cases and let it be judged on those four tests.
For the longer version of how it argues and how it is governed, read SIROC: an AI analyst that argues in the open — then book a demo and bring the ugliest case you have.