Skip to content

Startup checks

Three things can leave a sensor running perfectly while producing the wrong answer, and none of them shows up in its own output.

All three are checked once at startup and reported in the banner. All three are advisory: they warn, they never refuse a start, and --no-preflight (or preflight_checks: false) turns them off.

Nothing is listening on the OTLP endpoint

The exporter tolerates a collector that is down. That is correct at runtime — a receiver restarting mid-shift must not cost findings — but it means a sensor aimed at a dead port exports nothing, indefinitely, in silence.

One TCP connect at startup turns that into a line you can read:

Startup banner

  otlp endpoint   : http://127.0.0.1:4318  (NOTHING LISTENING)
  WARNING: nothing is listening on 127.0.0.1:4318
    the exporter tolerates a collector that is down, so every export to
    http://127.0.0.1:4318 will be discarded silently and no finding will
    reach it

Only a refused connection is reported as a problem. A name that will not resolve, or a connect that times out, is shown as unchecked — because neither proves anything is wrong, and a check that cries wolf gets disabled.

The runtime path is untouched. A collector that goes down after the sensor starts is still tolerated exactly as before.

Another sensor already writes this ledger

Two sensors at the same privilege level share a ledger directory and record every episode twice.

BEGIN IMMEDIATE keeps the hash chain intact, so the damage is logical duplication rather than corruption — which is worse in one respect, because nothing fails and the counts simply read high. A dashboard showing twice the real activity looks like a busy host, not a misconfiguration.

A pid-bearing sensor.owner file in the ledger directory makes the second one say so:

Startup banner

  WARNING: another ShadowClaw sensor (pid 61943) already writes
           /Users/you/Library/Application Support/ShadowClaw
    both will record the same activity, so every episode lands twice and
    counts in the ledger will read high

It is a marker, not a lock:

  • Nothing waits on it.
  • A start never depends on writing it.
  • An unclean exit leaves a file naming a dead pid, which the next start takes over silently.
  • Readers never claim it — --ledger watch tails ledger.jsonl and does not contend with the writer, so a viewer is never mistaken for a competing sensor.

Two sensors in different directories are not a conflict, so the root/per-user ledger split still separates a privileged run from an unprivileged one.

Endpoint Security will not admit this process

This is the worst of the three, because the sensor keeps running and reports clean on exactly the tactics it was enabled to catch.

macOS declines an Endpoint Security client that lacks Full Disk Access with an error that names neither Full Disk Access nor System Settings. So a two-second trial subscription to a single low-volume event translates it:

Startup banner

  endpoint sec.   : UNAVAILABLE -- Full Disk Access is not granted
  WARNING: Endpoint Security is enabled but unavailable -- Full Disk Access
           is not granted
    grant Full Disk Access to the binary that runs the sensor (the Python
    interpreter, or the installed launchd job)
    System Settings > Privacy & Security > Full Disk Access
    credential access, identity creation, privilege escalation and
    launch-item persistence will NOT be detected on this run
    agent-config persistence and public exfil surfaces still work -- they
    need no privilege

It warns only when Endpoint Security was actually asked for. Running without it is an ordinary configuration and does not deserve a warning on every start — that case gets the quieter coverage note below instead.

The same probe backs --self-test, so the answer is available before you commit to a deployment.

Degraded coverage is stated, not implied

A detector reporting clean because it was never able to look is indistinguishable, on a dashboard, from a host that is genuinely clean.

So an unprivileged sensor says which half it has:

Startup banner, no Endpoint Security

  endpoint sec.   : off (enable_esf is false)
  host-plane coverage is PARTIAL without Endpoint Security:
    visible  : agent-config persistence (by polling, so graded
               down), public exfil surfaces, lineage via ps
    NOT seen : credential reads, identity creation, privilege
               escalation, launch-item persistence, encoded payloads
    re-run under sudo with enable_esf to close the gap

The same concern drives shadowclaw.esf.running, which is exported on every poll and reads zero when the source is down. If it were emitted only while healthy, a subscription that died would leave no trace at all — and absence is the hardest thing to alert on.

That is worth an alert of its own. splunk/shadowclaw-detection.spl ships an "Endpoint Security went dark" search for exactly this.

What the banner always tells you

Beyond the three checks, every start reports the facts that determine what the run can possibly find:

Line Why it matters
Authorship and digest status A tampered identity block is announced, and telemetry is tagged attribution_intact=false.
config and its provenance Distinguishes a loaded policy from built-in defaults.
Privilege level Determines whether socket attribution covers one user or all of them.
Plane coverage 3 of 3 on macOS with root and FDA; fewer otherwise, with the reason.
otlp endpoint Plus the reachability result above.
Ledger directory Which of the root/per-user paths won.
Setting Default Effect
preflight_checks true All three checks above.
python3 -m shadowclaw --no-preflight

Turning them off is reasonable in a tight CI loop where the endpoint is known-dead by design. It is a poor idea on an endpoint.

Next