Metrics¶
Emitted every poll as OTLP gauges and monotonic counters. Enabled by emit_otlp_metrics, disabled by --no-otlp-metrics.
Loki does not accept these
Loki serves /otlp/v1/logs only and 404s the metrics path. Send metrics to a Collector or Prometheus. See Grafana and Loki.
Sensor health¶
| Metric | Kind | Meaning |
|---|---|---|
shadowclaw.sensor.up |
gauge | Liveness. |
shadowclaw.findings.active |
gauge | Findings in the current poll. |
shadowclaw.esf.running |
gauge | Emitted on every poll, including when zero. |
shadowclaw.esf.enabled |
gauge | Whether the host plane was asked for. |
shadowclaw.esf.events.received |
counter | Endpoint Security events parsed. |
shadowclaw.esf.events.dropped |
counter | Events lost to the bounded queue. |
shadowclaw.esf.queue.depth |
gauge | Current queue occupancy, ceiling 20,000. |
shadowclaw.configwatch.files |
gauge | Agent-config files being polled. |
esf.running is the one to alert on¶
It is exported on every poll and reads zero when the source is down. If it were emitted only while healthy, a subscription that died would leave no trace at all — and absence is the hardest thing to alert on.
An empty host-plane dashboard row is not a clean host. This metric is what tells the two apart, and splunk/shadowclaw-detection.spl ships an "Endpoint Security went dark" search for exactly this.
esf.events.dropped and queue.depth¶
A rising dropped counter means the reader is falling behind. eslogger open is a firehose — tens of thousands of events per second on an idle machine — which is why the credential stream runs behind a grep -F prefilter.
Measure before you deploy:
If the credential stream is not viable on your hardware, set esf_credential_stream: false. Exec-argv credential detection continues to work without it.
Detection¶
| Metric | Kind | Meaning |
|---|---|---|
shadowclaw.risk.score |
gauge | Per finding. |
shadowclaw.egress.connection |
gauge | Attributed egress connections. |
shadowclaw.process.cpu.percent |
gauge | Per implicated process, from CPU-time deltas. |
shadowclaw.process.memory.rss |
gauge | Per implicated process. |
shadowclaw.candidate.cpu.percent |
gauge | CPU for candidate runtimes not yet at threshold. |
candidate.cpu.percent is the leading indicator for inference_heartbeat: a process climbing here is one that has not yet exceeded cpu_percent_threshold for consecutive_samples polls.
Host plane¶
| Metric | Kind | Meaning |
|---|---|---|
shadowclaw.agent.sessions.active |
gauge | Agent sessions inside the chain window. |
shadowclaw.agent.tactic.observed |
counter | Tactic observations, by tactic. |
The count connector in config/otel-collector.agentic.yaml derives fleet-wide tactic rates from these, so a fleet trend question is answerable from the metrics store rather than a log scan.
Kinetic Trust Protocol¶
| Metric | Kind | Meaning |
|---|---|---|
ktp.risk_factor.adversarial_pressure |
gauge | Stress term, [0,1]. |
ktp.risk_factor.evidence_density |
gauge | Stress term, [0,1]. |
ktp.risk_factor.trust_trend |
gauge | Stress term, [0,1]. |
ktp.risk_factor.update_resistance |
gauge | Stress term, [0,1]. |
ktp.aggregation.silent_processes |
gauge | Tracked processes that stopped appearing. |
Each factor carries these attributes:
| Attribute | Meaning |
|---|---|
ktp.degraded |
Read this before reading the value. |
ktp.degraded_count, ktp.degraded_inputs |
Which inputs were substituted. |
ktp.feeds_active, ktp.feeds_total |
Coverage behind the value. |
ktp.confidence |
Confidence in the reading. |
ktp.risk_domain |
node |
ktp.spec.version |
2.0.0 |
ktp.source_id |
Emitting sensor identity. |
1.0 means hostile or unobserved
Anything the detector could not observe reports 1.0, never 0. On an unprivileged sensor adversarial_pressure never drops below 0.714.
Full explanation in Risk Factors.
Resource attributes¶
Every metric carries the same resource as every log record — product,
authorship, the attribution.intact official-build provenance indicator,
host.name, and a derived os.type. See
Event schemas. The indicator is product
metadata, not a license-compliance signal.
Host metrics from the Collector¶
config/otel-collector.yaml also runs the Collector's own hostmetrics receiver in two pipelines:
| Pipeline | Scraper |
|---|---|
metrics/ai-processes |
Process CPU, RSS, and thread counts. |
metrics/host |
Host-wide network and load. |
Those are host-wide — they tell you bytes moved, never which pid moved them. That is precisely why the sensor exists. See Why a sensor as well as the Collector.
Alert on findings rather than on raw host metrics: correlation happens at the endpoint where per-process attribution exists, and redoing it downstream is strictly weaker.