Security notes¶
The detector reads command lines, which is where API keys live. A monitoring control that mirrors credentials into a log index is worse than no control.
Redaction¶
Everything ShadowClaw emits passes through shadowclaw/redact.py before it is written or exported — including into the local ledger.
| Pattern | Examples |
|---|---|
| Provider key prefixes | sk-, sk-ant-, AIza, hf_, ghp_, AWS access key ids |
| Assignment shapes | KEY=value, --api-key value |
| Authorization headers | Bearer … |
| JWTs | Three base64 segments separated by dots |
Redaction is upstream of every output, so a cmdline in findings.jsonl, in the ledger payload, in Grafana, and in Splunk are all the same scrubbed string.
Transport¶
The exporter refuses to send plaintext to anything but loopback. This telemetry names the AI endpoints your staff reach and must not cross a routable interface unencrypted.
The OTLP receiver in the shipped Collector configs binds 127.0.0.1 only. Same reasoning.
File permissions¶
| Path | Mode |
|---|---|
| Ledger directory | 0700 |
ledger.db, ledger.jsonl |
0600 |
findings.jsonl |
0600 |
ktp-risk-factors.jsonl, ktp-envelope.jsonl |
0600 |
The ledger directory sits outside the repository, so cleaning the working tree cannot destroy the evidence — and make clean explicitly cannot reach it.
System-wide under root, per-user otherwise, because an unprivileged run's record is partial and would be misleading written to the machine-wide location. See The local ledger.
Secrets in configuration¶
Nothing sensitive is read from a config file.
| Secret | Source |
|---|---|
| Splunk HEC token | $SPLUNK_HEC_TOKEN |
| Splunk HEC URL | $SPLUNK_HEC_URL |
| Grafana API token | --token, or the environment |
config/otel-collector.yaml references the environment variables rather than holding values, so the file is safe to commit.
Subprocess invocation¶
Subprocesses are invoked with fixed argument lists, never a shell string. ps, lsof, tcpdump, eslogger and grep are all called with explicit argv arrays, so nothing observed on the host can reach a shell.
That matters more than usual here: the sensor's inputs are process names, command lines, and file paths chosen by whatever is running on the machine — including, potentially, an agent that would like to be misread.
Privilege¶
The sensor uses root for exactly two things:
| Capability | Why root |
|---|---|
| Full socket visibility | Unprivileged lsof returns established sockets for the invoking user only. |
| Endpoint Security | eslogger requires root, and Full Disk Access on many builds. |
Nothing else escalates. scripts/ShadowclawAI uses sudo for the sensor step only, and prints what it is doing.
Full Disk Access is granted to your terminal or IDE, not to ShadowClaw, and it lives in a SIP-protected database no script may edit. That means the uninstaller cannot revoke it — you have to, in System Settings → Privacy & Security, and only if nothing else you run needs it.
The Collector binary is not patched¶
scripts/install-otelcol.sh downloads the upstream OpenTelemetry Collector contrib build and verifies it against the upstream SHA-256.
It is deliberately not branded or modified. Modifying it would defeat the verification that makes the download trustworthy in the first place. See NOTICE.
Detection-only is a security property¶
ShadowClaw does not block, terminate, quarantine, or make exceptions, and there is no allowlist — with a test that fails the build if one is added back.
That is a security choice as much as a design one:
- A detector that can be told to stay quiet about a process no longer has a complete ledger, and its record stops being usable as evidence.
- The first thing anyone evading the control would do is get themselves added to the list.
- Blind spots that are configurable are blind spots you cannot reason about.
Exceptions belong to the enforcement tool downstream, applied with the finding in hand. See DefenseClaw.
Official-build provenance as tamper detection¶
The authorship headers and the SHA-256 digest in shadowclaw/authorship.py are
a deliberate, lightweight guard against a specific threat: an AI agent
editing this codebase and silently rewriting provenance, policy, or detection
logic as part of an otherwise plausible patch.
The headers make every touched file visible to tests/test_authorship.py. The digest makes any change to the canonical identity block observable at runtime and on every telemetry record.
This is simple tamper detection, not cryptographic proof. It makes a derivative's different identity metadata observable as distinct from the published build. It is not a licensing-compliance mechanism. See Authorship.
Reporting a security issue¶
Attribution and detection integrity are load-bearing in this project. If you find a way to make the sensor report a clean host while it is not, or to make a finding disappear from the ledger without breaking the chain, that is the most valuable class of bug here.