Graphic art for the blog titled "The Agentic SOC That Refuses to Guess" by D3 Security

The Agentic SOC That Refuses to Guess

TL;DR: A large US-based MSSP ran five adversarial tests against Morpheus APD 2.0’s AI SOC guardrails: broken Microsoft Sentinel authentication, a CrowdStrike agent outage, a query-by-query audit, cross-tenant probes, and Okta API drift. APD 2.0 passed all five. Zero false all-clear verdicts, a 100% query match, and tenant isolation in every search. When Morpheus is uncertain, it defers to a human.

Prospects keep asking us a version of the same question: how do I know the LLM won’t make a wrong decision? It’s the right question. An agentic SOC that turns incomplete data into a confident verdict is a liability with a good vocabulary.

What does the AI do when it can’t investigate? That is the hardest question in AI security operations.

Every vendor claims their AI SOC has guardrails. Almost none of those claims get tested the way one of our customers tested Morpheus Attack Path Discovery (APD) 2.0. They deliberately broke its data sources, degraded its telemetry, audited its every query, and probed for the one failure that would end a multi-tenant security business: data leaking between customers.

They ran five adversarial scenarios. APD 2.0 passed all five, with no critical findings. Here’s what they did.

The setup: a customer who refused to take our word for it

The customer is a large US-based MSSP operating a multi-tenant SOC with Microsoft Sentinel as a primary data source, CrowdStrike on the endpoint, and Okta for identity. They’ve asked us not to name them, and we won’t. Before putting APD 2.0 into service, their security engineering team designed their own guardrail acceptance tests. They wrote the scenarios, ran them in their own environment, and set the pass/fail criteria.

They targeted the failure modes that keep security leaders away from AI autonomy: the false all-clear, the unverifiable narrative, cross-tenant contamination, and silent degradation.

Scenario 1: Kill the SIEM connection

The team deliberately broke authentication to Sentinel, cutting APD 2.0 off from the SIEM mid-operation. The dangerous outcome here is silence. The AI queries, finds nothing because it can’t reach the data, and closes a real incident as benign.

That’s not what happened. APD 2.0 identified a permanent authentication issue and stated plainly that its verdict could not be confirmed. It produced no false negative and no manufactured confidence.

Scenario 2: Take the CrowdStrike agent offline

Same idea, different failure. The team made the CrowdStrike agent unavailable. APD 2.0 recognized a platform-wide issue and flagged in its investigation summary that results could not be fully confirmed.

Both scenarios prove the same property. APD 2.0 knows the difference between “no threat found” and “unable to verify.” When evidence is incomplete, it defers to human judgment. It fails safe, not silent.

Scenario 3: Audit every query

An investigation summary is only worth what it can prove. So the customer counted the Sentinel queries APD 2.0 executed and compared them against the queries listed in its investigation summary.

The counts matched. What APD 2.0 reports is what APD 2.0 does. For a CISO, that’s the property that makes autonomy governable. The investigation record is verifiable evidence your auditors can check, not a generated story you have to trust.

Scenario 4: Try to mix tenants

The customer’s most important test: could an investigation pull data across customer tenants? One investigation guideline required the customer name in every Sentinel search, and APD 2.0 honored it in every query it ran. Scoping applies at the field level of the SIEM data itself, so another tenant’s results cannot enter an investigation. The customer found no case in which an investigation could reach another tenant’s data.

The guideline is policy the customer wrote, and APD 2.0 enforced it in 100% of queries. When Morpheus is uncertain, it defers to a human. Guardrails you configure. Compliance you can measure.

Scenario 5: Degrade the telemetry, then watch the recovery

The fifth scenario shows what happens after a guardrail fires. API drift on Okta degraded the quality of identity telemetry feeding the investigation. APD 2.0 didn’t guess. It flagged that results could not be fully confirmed, and Morpheus routed the issue details to SOC analysts. After the fix, Morpheus reran the queries and reached the correct conclusion.

That’s the full loop security leaders should demand: detect the degradation, hand off to humans with the issue details, and re-verify once the problem is fixed. Guardrails are the handoff that keeps investigations accurate through real-world integration failures.

Why this matters

In the customer’s words:

“We did not observe any issues during the test. In all error scenarios, the investigation verdict was correct. Morpheus was able to avoid false negatives. Morpheus was able to provide the correct response from APD. More importantly, Morpheus failed to a human analyst when it had incomplete information.”

Security engineering team, large US-based MSSP

Autonomous investigation only earns a place in your SOC if it behaves correctly when things go wrong, and if you can verify that behavior yourself. This MSSP’s testing demonstrated four properties every security leader should demand from an agentic SOC.

It fails safe. Under induced data-source failures, APD 2.0 produced zero false “all-clear” verdicts. It reported the root cause and withheld conclusions it couldn’t stand behind. When Morpheus is uncertain, it defers to a human.

It’s provable. The audit trail matched reality, query for query.

It’s governed by your rules. Investigation guidelines acted as enforceable policy, down to tenant scoping in every search.

It recovers with humans in the loop. When telemetry degraded, Morpheus surfaced the issue to analysts and re-verified after the fix, then reached the correct conclusion.

That’s what we mean by the accountable agentic SOC platform: autonomy with governance at every stage and a traceable record of every action.

Preview of the case study titled The Agentic SOC That Refuses to Guess: A Large US-Based MSSP Tried to Break APD 2.0 and Couldn't by D3 Security

Want to hand this to your team? The case study is also available as a PDF, ready to forward or print

Don’t take our word for it either. The fastest way to trust an agentic SOC is to test one on your own terms, the way this customer did.

Book a Demo and run your own scenarios against APD 2.0

Related: APD 2.0 guardrails · MSSP case study · Agentic SOC glossary · Guardrails FAQ

Learn More About Morpheus

Powering the World’s Best SecOps Teams

Ready to see Morpheus?