Skip to content

Your first detection

A finding is not a single observation. It is a per-process rollup of every signal that fired for one pid inside one poll, with the endpoints and providers that justified them attached. This page reads one all the way through.

Generate something to catch

# Terminal 1
./scripts/run-sensor.sh --verbose

# Terminal 2
python3 simulators/rogue_agent.py --iterations 0

The simulator burns CPU in a scriptable runtime, reaches provider endpoints directly, and walks the host-plane tactics in order — so it exercises all three planes rather than only the easy one.

What the console prints

Withheld from the public documentation

Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.

Reading it in order:

[critical 100]
The risk score and its severity band. Signals are additive and the sum is capped at 100, so a finding with more corroboration than the ceiling allows saturates rather than running away. The bands, in order, are critical, high, medium, low and info; their cut-offs are in the signal catalog.
python3 [pid 41207] user alice
Process identity. python3 on macOS actually resolves to an executable named Python with a capital P — every regex in the detector is case-insensitive because of it, and there is a test pinning that.
shadow_ai_egress … via dns
A Plane B signal. The weight comes from the provider's category — 50 for a frontier LLM API. via dns is the attribution source, and it is the only exact one; catalog and ptr are weaker evidence an analyst should weigh differently.
gateway_bypass
Fires only because sanctioned_endpoints declares an approved gateway. This is the difference between "someone used AI" and "someone evaded the control". It fires even when the process also used the gateway — mixed traffic is not exculpatory.
inference_heartbeat
A Plane A signal: sustained CPU across consecutive polls in a scriptable runtime. Derived from CPU-time deltas, not ps %cpu, which is a lifetime average that would never show the burst.
agent_credential_access
A Plane C signal, and note the lineage: claude -> sh -> cat. No host-plane signal fires for a process without an AI agent above it in the process tree. The same cat run by a developer produces nothing.
agent_kill_chain
Three or more distinct tactics, in forward order, inside the session window. Progression is worth more than the sum of its parts: reading a credential is a lead; reading a credential then minting an identity then uploading to a paste site is an incident.

Withheld from the public documentation

Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.

Where the result lands

Output order is deliberate — the ledger is written before any file, console, or network export, so a Splunk or Collector outage can never cost you the record.

First

Durable local record

ledger.db in SQLite, plus the raw append-only ledger.jsonl. Hash-chained and immutable, outside the repository so make clean cannot destroy it.

  • events — one row per observation
  • episodes — mutable rollup
  • --ledger verify checks the chain

Then

Local files and console

findings.jsonl beside the ledger, and the console line above. Both are optional — --no-file, console: false.

  • 0600 inside a 0700 directory
  • Relative paths anchor to the ledger dir
  • "" switches the stream off

Last

OpenTelemetry export

Five flat log event shapes plus gauges and counters, posted as OTLP/HTTP JSON with no SDK in the path.

  • finding.recorded
  • provider.reached, agent.activity
  • ktp.risk_factors, ktp.envelope

The same finding as JSON

findings.jsonl and the ledger carry the full structure. Abridged:

{
  "finding_id": "9f2c41ab7d0e5613",
  "pid": 41207,
  "process_name": "python3",
  "exe_path": "/usr/bin/python3",
  "user": "alice",
  "risk_score": 100,
  "severity": "critical",
  "signals": [
    {
      "signal_id": "shadow_ai_egress",
      "title": "Direct connection to a known AI provider",
      "weight": 50,
      "detail": "api.anthropic.com (Anthropic) via dns attribution"
    }
  ],
  "endpoints": [
    {
      "remote_ip": "203.0.113.10",
      "remote_port": 443,
      "hostname": "api.anthropic.com",
      "attribution_source": "dns",
      "provider_id": "anthropic",
      "category": "frontier_llm_api",
      "sanctioned": false,
      "confidence": 0.95
    }
  ],
  "cpu_percent": 43.2,
  "rss_mb": 512.0,
  "detected_by": {
    "product": "ShadowClaw",
    "version": "1.5.1",
    "attribution_intact": true
  }
}

Command lines are scrubbed by shadowclaw/redact.py before anything is written. See Security notes.

Next