Skip to content

Signal catalog

Twenty-one signals across three planes. Weights are additive, the per-process sum is capped at 100, and only findings at or above the configured reporting floor, min_risk_to_report, are emitted.

Withheld from the public documentation

Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.

Severity bands

Withheld from the public documentation

Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.

Planes A and B

Withheld from the public documentation

Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.

Notes on the variable ones

shadow_ai_egress
Weight is the provider's category weight: 50 for a frontier LLM API or an inference aggregator, 35 for a cloud AI platform, 30 for a consumer chat UI or model hub, 25 for a coding assistant or agent-observability service.
gateway_bypass
Only possible when sanctioned_endpoints is populated. Fires when a process reaches a provider directly while that gateway is declared — including when the process also used the gateway. Mixed traffic is not exculpatory.
unattributed_tls_egress
Starts below the reporting floor on its own, and escalates once the same scriptable runtime reaches several distinct unattributed peers. Repetition is the signal, and this is the backstop against DNS-over-HTTPS.
sanctioned_ai_egress
Deliberately below the reporting floor. It exists so approved use appears in the inventory and in the ledger rather than being hidden — the difference between labelling and muting.

Plane C — host

Every one of these requires an AI agent in the process lineage. See Plane C.

Withheld from the public documentation

Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.

agent_encoded_payload is corroboration-gated. Encoding by itself is what ordinary build scripts do, and a base64 pipeline cannot reveal whether its input was a private key or generated test data. It contributes neither weight nor a kill-chain stage until the same session also supplies evidence of something worth packaging (credential access) or somewhere to send it (a catalogued/shared exfiltration surface). The observation is retained while it waits, so packaging that occurs before the upload is not lost.

Transfer-shaped commands are different evidence and use agent_payload_transfer: curl --upload-file, curl @-, and netcat to a remote endpoint remain ungated. Splitting the signal is what prevents a base64-noise fix from silencing an upload to an uncatalogued host.

Withheld from the public documentation

Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.

The two shared-channel signals

Unlike every other Plane C signal, both of these are evidence about a relationship between two agent sessions rather than a stage either session took alone — motivated by OpenAI's disclosed Hugging Face incident and its separate, recurring "Artifactory as message board" misalignment reports. Neither participates in agent_kill_chain's ordered progression for that reason; see Agent kill chain.

agent_shared_channel_use_local

One session writes to a path outside any git working tree; a different session later reads it. Full confidence outside any known package-manager cache root, graded to 0.6 under one (~/.cache, ~/.npm, ~/.cargo/registry, ~/.gem, ~/.m2, ~/.nuget, Go's module cache) — two unrelated agents both touching a shared build cache is weaker evidence than two sessions touching an unexplained path.

The finding is reported against the reader, not the writer, and names the writer inside it (written by cursor[38441] 25s earlier). A write on its own is indistinguishable from any other file being written; the correlation only becomes knowable when the read arrives, so that is the event there is to report. Expect one timeline entry per detection, not two.

Two acquisition requirements, both of which produced a silently dead signal before they were understood:

A read is only matched to a write that already happened. Events from the signal, credential and write subscriptions arrive on one queue in no particular order, so without an explicit ordering check the correlation would report writes from the future — observed, before the check existed, as written by codex[66623] -68s earlier.

agent_shared_channel_use_external

The same pattern against a domain already in the public exfiltration-surface catalog — paste sites and dead-drop services, once an agent has internet access. Fixed at confidence 0.5 regardless of surface kind: over TLS this can only confirm both sessions reached the same domain, not the same paste or file, which is materially weaker than the identical-path case above. At weight 35 this scores 18 — below the default reporting floor alone, so it corroborates without ever alerting by itself.

Brief contact is observed from the streamed PKTAP connection source rather than inferred from the lsof sample: outbound TCP SYNs and QUIC/HTTP-3 Initial packets carry the process identity at connection initiation. lsof remains additive for duration and settled socket state.

A surface sharing addresses with a larger parent domain no longer disappears behind the generic winner. The resolver retains every observed DNS/catalog candidate for an IP and the exfiltration path inspects only catalogued transfer surfaces among them. Generic provider scoring still receives one deterministic attribution, so fixing the surface case does not guess a provider.

When a transfer surface and an unrelated parent/CDN name both remain viable, the surface evidence is marked ambiguous and discounted. It can corroborate other behavior but cannot establish that two sessions used the same external channel; ordinary traffic to a shared GitHub edge must not become a fictitious Gist handoff.

A gateway that intercepts the connection before it completes can still leave no destination flow to attribute, though in that case the egress was also prevented.

Signal-to-tactic mapping

Eight host signals map onto the six ordered tactics used by the kill chain. Two more map onto their own standalone tactics, deliberately excluded from that ordering — see the note above.

Signal Tactic ATT&CK
agent_local_mcp_server local_inference T1059
agent_credential_access credential_access T1552
agent_identity_creation identity_creation T1136
agent_privilege_escalation privilege_escalation T1548
agent_persistence persistence T1543
agent_config_persistence persistence T1543
agent_encoded_payload exfiltration T1567
agent_payload_transfer exfiltration T1567
agent_public_exfil_surface exfiltration T1567
agent_shared_channel_use_local inter_agent_channel_local (standalone) T1559
agent_shared_channel_use_external inter_agent_channel_external (standalone) T1102

ATT&CK enrichment — technique URLs and tactic names — happens in the Collector, not the sensor, because those mappings are the part most likely to need correcting, and correcting them there does not mean shipping a new sensor to every endpoint.

Confidence discounts weight

Attribution and event confidence are not binary. A signal supported by evidence with confidence below 0.9 contributes proportionally less than its nominal weight.

That is why via dns and via catalog produce different scores for the same hostname, and why a webhook_relay (confidence 0.6) scores lower than a paste_service (1.0) for the same agent_public_exfil_surface signal. The two shared-channel signals apply the same rule: agent_shared_channel_use_local at 0.6 under a package-manager cache root, and agent_shared_channel_use_external fixed at 0.5 for its inherently weaker, domain-only correlation.

Every path is exercised

python3 -m shadowclaw --verify

Runs 37 synthetic scenarios against the real scoring code — 13 of them benign cases that must stay quiet, including a developer running sudo, a legitimate installer writing a LaunchDaemon, an MDM-managed launch item, base64 inside a build, an agent reading its own MCP config, and two agents sharing one project's own git tree without being mistaken for a covert channel. See Verification.

Next

Agent kill chainHow the +25 bonus is earned.

Provider catalogWhich providers carry which category weight.

ConfigurationThresholds, windows, and the reporting floor.