Skip to content

Verification

Watching a detector light up is not evidence. You notice the providers it caught and not the ones it missed. So there are three separate things to check, and they check different claims.

Detection matrix37 scenarios
Benign cases13
Unit tests2,196
Test files51

The detection matrix

python3 -m shadowclaw --verify
# or
make verify

Runs 37 synthetic scenarios through the real scoring, tactics, and chain code — no network, no privilege, no host state. Deterministic, so it is safe in CI.

python3 -m shadowclaw --verify

  [PASS] agent-deep-descendant            risk=30 medium [agent_credential_access]
         Lineage survives three levels of shell between agent and action
  [PASS] benign-developer-sudo            no finding, as expected
         A developer running sudo has no agent behind it and scores nothing
  [PASS] benign-installer-launch-daemon   no finding, as expected
         A legitimate installer writing a LaunchDaemon is not an agent
  [PASS] benign-mdm-managed-launch-item   no finding, as expected
         An MDM-managed launch item is expected, even under an agent
  [PASS] benign-base64-in-build           no finding, as expected
         base64 under an agent is the most ordinary command there is
  [PASS] benign-agent-reads-own-config    no finding, as expected
         Reading its own MCP config is what an agent is supposed to do
  ...
------------------------------------------------------------------------------
  37/37 scenarios passed

Thirteen of them must stay quiet

That is the half most detection test suites skip, and it is the half that decides whether the tool survives contact with a real developer machine.

Benign case What it pins
A developer running sudo Lineage gating. No agent above it, no signal.
A legitimate installer writing a LaunchDaemon Persistence needs an agent, not just a launch item.
An MDM-managed launch item, even under an agent Managed configuration is expected.
base64 inside a build The most ordinary command there is.
An agent reading its own MCP config Reading is what an agent is supposed to do.
A polled config change with no agent running Attributed to nobody, so it accuses nobody.
A polled config change on its own Circumstantial evidence corroborates a chain; it does not raise one.
A compiler touching a shell rc file No agent, no signal.

A regression that makes the detector more sensitive fails this suite exactly as loudly as one that makes it blind.

The coordinated test

The matrix proves the scoring is correct. It does not prove the acquisition works on your host, against real traffic, at your sample interval.

So declare the ground truth first and let the run be scored against that claim:

scripts/coordinated-test.py --expect anthropic,fireworks --duration 90

It records the window, runs the sensor blind to your list, then scores the durable ledger against it.

Column Meaning
DETECTED You declared it and it was found.
MISSED You declared it and it was not.
SHADOW Found, but not declared. Reported separately, not counted against you.

That third column is usually the interesting one, because it is the real AI activity already running on the host as opposed to the traffic you generated on purpose.

Accepts ids, display names, domains, or aliases — claude resolves to Anthropic, gemini to Google.

Drive the traffic yourself when prompted, or hand it over:

scripts/coordinated-test.py --expect anthropic,fireworks,google_ai \
  --swarm "python3 simulators/rogue_agent.py --iterations 0" \
  --report scorecard.txt

Exit status is 0 only when every declared provider was detected, so it can gate a build.

make coordinated-test EXPECT=openai,groq

Endpoint Security volume

Before trusting the host plane in the hot path, measure what it will actually cost on your hardware:

sudo python3 scripts/esf-volume-spike.py

The tool measures rather than estimates: raw and prefiltered open, create/write/rename, parser throughput, and program-axis selection over paired windows. On the development host unfiltered open was a real but acceptable single-digit share of one core, not the “tens of thousands per second” the module once claimed without measurement.

Unfiltered is the default because a path prefilter makes its fragment list the hard boundary of what file activity can be detected. The stream is bounded and counts overflow, so a busier host degrades visibly instead of silently falling behind.

Event-driven connection coverage

Compare streamed PKTAP initiation events against the additive lsof snapshot over one shared wall-clock window:

sudo python3 scripts/pktap-vs-lsof.py --seconds 120

The report separates pre-existing sockets, inbound connections, non-QUIC UDP, and genuine misses rather than publishing one misleading “lsof only” total. PKTAP can lead only when genuine misses are zero; lsof still remains for settled state and duration.

Windows argv acquisition

Command-line-derived Plane C signals require argv in the process-creation event. A process-table query is not an acceptable fallback because short-lived children can exit before the query:

py -3 packaging\windows\tests\test-windows-argv-acquisition.py

Run elevated after enabling Process Creation auditing and command-line inclusion, as documented by the script. It launches one fixed-marker child and requires both the private SystemTraceProvider stream and Security 4688 to carry that marker. The script never prints arbitrary captured argv.

The unit suite

make test

2,196 tests across 51 files, with no runtime dependencies. Beyond ordinary coverage, several tests enforce structural properties that are easy to erode:

Test Fails the build if…
test_ktp_zero_dependency.py Anything under shadowclaw/ imports a third-party module. Checked by walking the source and by importing the sensor in a fresh interpreter.
test_authorship.py The identity constants change, any source file loses its header, or telemetry stops carrying the author.
test_tactics_seam.py tactics.py, scoring.py or agentchain.py names an operating system or asks which one it is on.
test_platforms.py The macOS indicator set is no longer object-identical to the tactics constants.
test_ktp_envelope.py A detection module so much as imports the Kinetic Envelope.
(policy test) An allowlist is added back.
test_ktp_envelope_properties.py The Envelope's six declared properties fail over the input grid.

test_a_dark_host_plane_does_not_read_as_quiet is the pin on the single most important safety property: a sensor that cannot see is never reported as a calm host.

Collector config validation

make validate

Validates config/otel-collector.yaml against the installed Collector binary — so a pipeline typo is caught before a deployment quietly drops telemetry. Requires ./scripts/install-otelcol.sh first.

What none of this proves

Verification confirms the detector behaves as designed on the evidence it is given. It does not prove an endpoint is free of unsanctioned AI use.

A determined insider can rename a binary, tunnel over an allowed host, use a browser extension, or break the process tree with cron. This is a governance and hygiene control, not an exfiltration-proof boundary — see Known limitations.

Next