Event schemas¶
Five log event shapes. All of them flat.
Why everything is flat
No nested objects and no arrays — evidence included — because consumers parse the line with | json (LogQL) or spath (SPL), and a nested attribute is neither a label nor a parseable field.
Multi-valued fields are comma-joined strings. LogQL and SPL match them with a regex; an array would not survive the parse. It is also why counting providers correctly has to happen at emit time rather than in the query.
Always filter on event_name
All five shapes carry overlapping process fields. A query that does not discriminate will silently mix a scored finding with the individual observations behind it, and double-count.
shadowclaw.finding.recorded¶
One record per scored finding.
| Field | Type | Meaning |
|---|---|---|
event_name |
string | shadowclaw.finding.recorded |
finding_id |
string | 16 hex characters. |
pid, process_name, exe_path, cmdline, user |
Process identity. cmdline is redacted. |
|
risk_score |
int | 0–100, capped. |
severity |
string | critical | high | medium | low | info |
summary |
string | Human-readable one-liner. |
signals |
joined string | Signal ids that fired. |
providers |
joined string | Every provider reached. |
endpoints |
joined string | Every endpoint touched. |
categories |
joined string | Provider categories. |
provider |
string | The single headline provider. See the warning below. |
category, sanctioned, confidence, attribution |
Attributes of the headline endpoint. | |
cpu_percent, rss_mb |
float | Resource use at detection. |
tactics, techniques |
joined string | Present when the host plane contributed. |
Do not group by provider
The headline provider is chosen by ranking endpoints: unsanctioned egress first, then a named provider, then confidence, then hostname alphabetically. One process talking to two providers produces one finding, so grouping by provider counts only the winner — and because it is a sort rather than a race, the loser is invisible permanently.
Use shadowclaw.provider.reached instead. See why.
shadowclaw.provider.reached¶
One record per distinct provider a finding reached — not per endpoint, so a provider behind four rotating CDN addresses still counts once.
| Field | Meaning |
|---|---|
event_name |
shadowclaw.provider.reached |
finding_id |
Links back to the finding. |
provider, category, sanctioned |
The provider reached. |
confidence, attribution |
How the attribution was made. |
endpoint |
The hostname or address. |
is_headline |
true for the one that would have been the finding's provider. |
severity, risk_score |
Carried from the finding, so a severity filter selects the same population as on the findings panels. |
pid, process_name, user |
Process identity. |
Filtering to is_headline="false" shows exactly what a finding-grouped panel drops.
Unattributed egress emits no provider record, so charts that need it read finding.recorded with provider = "".
shadowclaw.agent.activity¶
One record per newly observed tactic. This is the timeline an incident is reconstructed from.
{
"event_name": "shadowclaw.agent.activity",
"agent_name": "claude",
"agent_root_pid": 1000,
"attribution_depth": 2,
"attribution_authority": 0.01,
"attribution_state": "attributed",
"tactic": "credential_access",
"technique": "T1552",
"chain_stage": 1,
"signal": "agent_credential_access",
"detail": "opened /Users/~/.aws/credentials",
"confidence": 1.0,
"pid": 1002,
"process_name": "cat",
"source_event": "open",
"file_path": "/Users/~/.aws/credentials"
}
| Field | Meaning |
|---|---|
agent_name, agent_root_pid |
The session root. |
attribution_depth |
Process-tree levels between agent and action. |
attribution_authority |
Strength of the claim at that depth. |
attribution_state |
attributed, or the reason it is not. |
tactic |
One of the six ordered tactics. |
technique |
ATT&CK id. |
chain_stage |
Position in the ordered chain, 0–5. |
signal |
Which host-plane signal produced this. |
confidence |
Event confidence, before weighting. |
source_event |
The underlying Endpoint Security event. |
file_path |
Present for file-based tactics. |
Empty without the host plane
Requires --esf under root. A stream with no agent.activity records means the host plane is off, not that the host is clean. shadowclaw.esf.running is the authoritative answer to which.
shadowclaw.ktp.risk_factors¶
One record per poll — including quiet polls, so a gap in this stream means the sensor stopped, not that nothing happened.
{
"event_name": "shadowclaw.ktp.risk_factors",
"summary": "2 of 4 inputs unobserved (adversarial_pressure,trust_trend) -- values include maximum-stress substitutes, not measurements",
"degraded_inputs": "adversarial_pressure,trust_trend",
"degraded_count": 2,
"inputs_total": 4,
"fully_observed": false,
"feeds_active": 8,
"feeds_total": 11,
"feed_coverage": 0.7273,
"silent_processes": 0,
"adversarial_pressure": 0.714286,
"evidence_density": 0.08,
"trust_trend": 1.0,
"update_resistance": 0.17
}
Never read a factor value without degraded_inputs or degraded_count beside it. See Risk Factors.
shadowclaw.ktp.envelope¶
One record per attributed agent action.
| Field | Meaning |
|---|---|
autonomy_demand |
A, what this action reaches. |
environmental_capacity |
E, how well it can be seen right now. |
margin |
1 - A/E. Floor of −1.0. |
supervision |
stable | metacognitive | assisted | regulated | silent_veto |
capacity_known |
Read this before reading supervision. |
clamped |
true if the level was raised because capacity was unknown. |
vetoed |
true if demand met or exceeded capacity. |
evidence |
Comma-joined evidence codes, e.g. KINETIC_CAPACITY_EXCEEDED. |
profile |
shadowclaw-software-agent-v1@1 |
See Kinetic Envelope.
Resource attributes¶
Carried on every log record and every metric.
| Attribute | Value |
|---|---|
service.name, service.namespace |
shadowclaw, both. |
shadowclaw.product |
Product name. |
shadowclaw.author |
Mike Storm |
shadowclaw.author.title, shadowclaw.author.credential |
Title and credential. |
shadowclaw.copyright |
Copyright notice. |
shadowclaw.attribution.intact |
Official-build provenance indicator; false if the identity digest no longer matches. |
host.name |
Host identity. |
os.type |
Derived from the platform layer. |
On resource rather than record attributes so official-build provenance survives being forwarded. This is product metadata, not a license-compliance signal.
The internal finding¶
findings.jsonl and the ledger payload column carry the full nested structure — this is the one place nesting is allowed, because nothing queries it with LogQL.
{
"finding_id": "9f2c41ab7d0e5613",
"pid": 41207,
"process_name": "python3",
"exe_path": "/usr/bin/python3",
"cmdline": "python3 agent.py",
"user": "alice",
"risk_score": 100,
"severity": "critical",
"summary": "ShadowClaw (critical): python3 [pid 41207] -> api.anthropic.com",
"signals": [
{
"signal_id": "shadow_ai_egress",
"title": "Direct connection to a known AI provider",
"weight": 50,
"detail": "api.anthropic.com (Anthropic) via dns attribution"
}
],
"endpoints": [
{
"remote_ip": "203.0.113.10",
"remote_port": 443,
"hostname": "api.anthropic.com",
"attribution_source": "dns",
"provider_id": "anthropic",
"provider_name": "Anthropic",
"category": "frontier_llm_api",
"sanctioned": false,
"scope": "public",
"confidence": 0.95
}
],
"cpu_percent": 43.2,
"rss_mb": 512.0,
"first_seen": 1787210000.0,
"last_seen": 1787210010.0,
"detected_at": "2026-08-20T05:40:00-0700",
"detected_by": {
"product": "ShadowClaw",
"version": "1.5.1",
"author": "Mike Storm, Distinguished Engineer, CCIE Security 13847",
"attribution_intact": true
}
}
Ledger tables¶
| Table | Mutability | Contents |
|---|---|---|
meta |
write-once | Product, version, author, schema version, attribution digest, genesis and checkpoint hashes. |
events |
immutable, hash-chained | One row per observation, with prev_hash and hash. |
episodes |
mutable, not chained | Per-session rollup. A convenience, rebuildable from events. |
Only the immutable log is evidence. See The local ledger.