Agent kill chain¶
Five separate per-process findings would be five alerts nobody joins up. shadowclaw/agentchain.py keeps process ancestry, attributes every observation to the agent ultimately responsible for it, and scores the accumulated chain against one AgentSession.
The six tactics, in order¶
Order matters. The chain bonus requires distinct tactics observed in forward sequence, not merely present together.
flowchart LR
T1["local_inference<br/><small>T1059</small>"] --> T2["credential_access<br/><small>T1552</small>"]
T2 --> T3["identity_creation<br/><small>T1136</small>"]
T3 --> T4["privilege_escalation<br/><small>T1548</small>"]
T4 --> T5["persistence<br/><small>T1543</small>"]
T5 --> T6["exfiltration<br/><small>T1567</small>"]
| Stage | Tactic | ATT&CK technique |
|---|---|---|
| 0 | local_inference |
T1059 — Command and Scripting Interpreter |
| 1 | credential_access |
T1552 — Unsecured Credentials |
| 2 | identity_creation |
T1136 — Create Account |
| 3 | privilege_escalation |
T1548 — Abuse Elevation Control Mechanism |
| 4 | persistence |
T1543 — Create or Modify System Process |
| 5 | exfiltration |
T1567 — Exfiltration Over Web Service |
Lineage attribution¶
Every host-plane observation is walked back up the process tree until an AI agent is found. If none is found, nothing is reported — see Plane C.
claude (pid 1000)
└─ sh (pid 1001)
└─ sh (pid 1002)
└─ cat ~/.aws/credentials (pid 1003)
→ attributed to: claude, depth 3
The synthetic detection matrix includes a case named agent-deep-descendant that pins this: lineage survives three levels of shell between the agent and the action.
Attribution authority¶
Depth is not free. An observation three shells below the agent is a weaker claim about that agent's intent than one made directly, and the finding says so. Each activity record carries:
| Field | Meaning |
|---|---|
agent_name |
The recognised agent — claude, codex, cursor, aider, … |
agent_root_pid |
The pid the session is rooted at. |
attribution_depth |
How many levels of process tree separate agent from action. |
attribution_authority |
How strong the claim is at that depth. |
attribution_state |
attributed, or the reason it is not. |
Where lineage comes from¶
| Source | Quality |
|---|---|
Endpoint Security exec / fork / exit events |
Authoritative. Catches the eight-millisecond cat that polling never sees. |
ps parent pids |
Fallback when unprivileged. Misses anything shorter than the poll interval. |
Pids are reused. Lineage is rebuilt from exec events and re-parented on reuse, but a chain assembled across a recycled pid on a very busy host can in principle attribute an action to the wrong session.
Recognised agents¶
Session roots are recognised from process names and command-line shapes: Claude, Codex, Cursor, Aider, Goose, Windsurf, Copilot, Devin, OpenHands, SWE-Agent, AutoGPT, and CrewAI-style names, plus generic agent-framework and MCP command patterns.
The agent catalog is a list, and lists are incomplete
A tool nobody has heard of yet, or one deliberately renamed, roots no session. Behavioural signals on planes A and B still fire; the host plane goes quiet. This is the acknowledged cost of lineage gating.
Withheld from the public documentation
Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.
Withheld from the public documentation
Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.
The activity stream¶
Findings tell you that an agent misbehaved. The activity stream tells you in what order, and it is a separate event type for exactly that reason.
shadowclaw.agent.activity emits one record per newly observed tactic:
{
"event_name": "shadowclaw.agent.activity",
"agent_name": "claude",
"agent_root_pid": 1000,
"attribution_depth": 2,
"attribution_state": "attributed",
"tactic": "credential_access",
"technique": "T1552",
"chain_stage": 1,
"signal": "agent_credential_access",
"detail": "opened /Users/~/.aws/credentials",
"confidence": 1.0,
"pid": 1002,
"process_name": "cat",
"source_event": "open",
"file_path": "/Users/~/.aws/credentials"
}
That is the timeline you reconstruct an incident from. Full schema in Event schemas.
Reading a chain back¶
Two shipped paths reconstruct a session tactic by tactic:
splunk/shadowclaw-detection.spl includes an agent kill chain search and an incident timeline search that walks one agent session stage by stage. See Splunk.
The bundled dashboard carries a Host plane row: kill chains, credential reads, identities created, persistence, exfil surfaces, tactics over time by stage, and the per-tactic activity timeline. See DefenseClaw integration.
An empty Host plane row is not a clean host
Those panels render empty unless the sensor runs with --esf under root. Worth knowing before reading emptiness as safety.
The collector pipeline¶
Host-plane telemetry has its own complete Collector config rather than edits to the base one:
Same reasoning as the Splunk split: host-plane detection is opt-in, and turning it on must not be able to regress a working shadow-AI deployment. It adds the logs/agentic pipeline, derives technique URLs and ATT&CK tactic names in a transform processor, and runs a count connector so fleet-wide tactic rates are answerable from the metrics store rather than a log scan.