D3 Security · Security Operations Glossary
What Is Agent Washing?
A standalone glossary definition, part of the D3 Security Operations Glossary.
Definition
Agent washing is marketing that presents a general-purpose model summarizing alerts as autonomous investigation. The test is the record: an investigation can show the queries it ran and the evidence it weighed; a summary cannot.
The term exists because of how this category grew. In the first wave of AI security operations products, having an AI story was a market requirement, and for most vendors the story was the same: take the alarm, pass it through a general-purpose model, and return a better-written summary of what the alarm already said. That is narrative dressing on an unchanged queue. The analyst still reads every alert, still pulls the process list by hand, still queries the identity provider. Only the prose improved.
Where the term comes from
Gartner introduced the phrase in a June 2025 analysis of agentic AI, defining agent washing as the rebranding of existing products, including AI assistants, robotic process automation, and chatbots, without substantial agentic capabilities. In the same analysis the firm estimated that only about 130 of the thousands of vendors claiming agentic AI were real, and predicted that more than 40% of agentic AI projects would be canceled by the end of 2027 on escalating costs, unclear business value, or inadequate risk controls.
Security operations inherited the general problem in a specific form. In the SOC, the rebranded product is usually a summarizer. When every vendor in a briefing claims the same three things, the claims carry no information, and the words that were supposed to signal capability have stopped doing any work.
The record is the test
There is one reliable way to separate a summarizer from an investigator, and it does not involve reading the datasheet. Ask the system to show its work on your own data.
A summarizer produces better prose about the alert it was handed. An investigation produces the queries it ran, the evidence it retrieved, the case for escalation and the case for closure, and a record connecting evidence to conclusion to action. If that record does not exist, the autonomy is a claim, not a capability.
| What you ask for | A summarizer produces | An investigation produces |
|---|---|---|
| The alert write-up | Cleaner prose about the alert it was handed. | The same write-up, plus the reasoning that produced it. |
| Queries against your data | None. The model sees only the alert payload. | The exact queries run, which systems they hit, and when. |
| Evidence retrieved | The alert fields it was given. | What came back, what came back empty, and what stayed inconclusive. |
| The opposing case | A single verdict. | The case for escalation and the case for closure, both argued. |
| Behavior when evidence is missing | Confidence unchanged. The summary still reads well. | Confidence drops and the alert goes to a human. |
| Chain of custody | Not available. | Intake to action, including who approved what and under whose authority. |
This is why a scripted demo settles nothing. A scripted demo shows a path the vendor chose. The demonstration that matters is the system working on your alerts, in your environment, with the record available for inspection afterward.
Also see:
Triage Slop
AI Alert Triage
The vocabulary that stopped meaning anything
Three words survived the first wave and calcified. Each now covers so much ground that it tells a buyer almost nothing:
- Autonomous: describes everything from closed-loop response to drafting an email for a human to send.
- Human in the loop: means a genuine approval gate at one vendor and a rubber stamp after the fact at another.
- AI-powered: attaches equally to a purpose-built reasoning system and to a general-purpose model rewriting an alert.
Better adjectives will not settle any of this. Behavior will. Every claim in the list above can be tested by asking the vendor to demonstrate it on a real alert and then produce the trail.
Claims to test, and the evidence to demand
Four claims come up in almost every briefing. Each one has a specific artifact that settles whether the capability is real:
- Completely autonomous, no analysts required: ask for the record of a wrong decision, how it was caught, what it cost, and what changed afterward.
- AI-powered triage: ask for a live investigation on your data, with the queries run, the evidence retrieved, and both cases argued.
- Human in the loop: ask for the exact boundary, meaning what the system cannot do without approval and how that boundary is set.
- Simple per-alert pricing: ask for the bill at ten times your volume, at unsubsidized model rates, with retries and failed loops priced in.
Apply the same demands to every vendor on the shortlist, including D3. A vendor who cannot produce these artifacts has told you something useful about the product.
What an accountable agentic SOC shows instead
In Morpheus AI, each of those requests returns an artifact. Read-only investigation runs against your systems through Attack Path Discovery and fails toward a human when it cannot get what it needs. Scoring opens to its factors, its weights, and the evidence behind each one, including the evidence that cut against the verdict. A chain of custody runs from alert intake to final action, recording what was read, what was queried, what came back, what was concluded, what was recommended, what was executed, and under whose authority.
Weak evidence never closes an alert. When Morpheus is uncertain, it defers to a human. Both properties are commitments a buyer can check during an evaluation.
Frequently asked questions
What is agent washing?
Marketing that presents a general-purpose model summarizing alerts as autonomous investigation. The label describes the gap between what the copy claims and what the system does.
Who coined the term agent washing?
Gartner, in a June 2025 analysis of agentic AI. The firm defined it as the rebranding of existing products such as AI assistants, robotic process automation, and chatbots without substantial agentic capabilities, and estimated that only about 130 of the thousands of vendors claiming agentic AI were real.
How do you tell a real agentic SOC from agent washing?
Ask the system to show its work on your data. A summarizer produces better prose about the alert it was handed. An investigation produces the queries it ran, the evidence it retrieved, the case for escalation and the case for closure, and a record connecting evidence to conclusion to action. If the record does not exist, the autonomy is a claim, not a capability.
Why is a scripted demo insufficient?
A scripted demo shows a path the vendor chose in an environment the vendor controls. It cannot tell you how the system behaves on your alert mix, your log shapes, or your query languages, and it cannot show you what happens when evidence is missing.
Does agent washing mean the vendor is lying?
Not necessarily. Much of it comes from a category where the vocabulary got ahead of the systems, so words like autonomous and human in the loop are used sincerely to describe very different things. The remedy is the same either way: ask for the artifact, not the adjective.
What single question exposes it fastest?
Show me the complete record of one real triage decision: what was read, what queries ran, what evidence returned, what was concluded and why, what executed, and who approved it. Treat the absence of that record as a finding.
Is summarizing alerts useless?
No. A well-written summary saves reading time. The problem is pricing and positioning a summary as investigation, because the buyer then expects analyst hours back and does not get them.
How does missing evidence expose agent washing?
Disable a log source or revoke a credential during the evaluation and watch what the system concludes. A system that grows more confident as evidence disappears is not performing investigation, whatever the marketing says.
Should the same tests apply to D3?
Yes. Every demand in this entry applies to Morpheus AI on the same terms as to any other platform on your shortlist.
Related terms
Triage Slop — Low-quality automated triage output that a black-box score can hide.
AI Alert Triage — Automated investigation and disposition of alerts at machine speed.
Governed Agentic SOC — The operating model that keeps autonomous triage inside explicit governance.
Agentic SOC — A security operations model in which AI agents autonomously triage, investigate, and respond to alerts while human analysts supervise.
Cybersecurity Triage Reasoning Graph — The reasoning engine that carries the judgment behind the score.
Further reading
The SOC After the Agentic SOC
Why fail-open matters
Attack Path Discovery
Book a demo
Last updated: July 2026