Plane A — inference heartbeat¶
Plane A answers is this process doing AI? using only what the machine is computing. It needs no catalog, no network attribution, and no privilege.
Why CPU time, not ps %cpu¶
ps %cpu is a lifetime average, not a current reading. A process that has been idle for an hour and then spends thirty seconds running inference shows a %cpu in the low single digits — it will never show the burst you are looking for.
The sensor differences cumulative CPU time between polls instead:
which is why a spike appears within one sample interval rather than being averaged into invisibility. This is one of two macOS details that silently break naive implementations; the other is that python3 resolves to an executable named Python with a capital P, so every process regex in the detector is case-insensitive and a test pins it.
Signals¶
Withheld from the public documentation
Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.
Thresholds¶
All three live under heartbeat in the config.
Withheld from the public documentation
Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.
What counts as a scriptable runtime¶
inference_heartbeat only fires for processes that could plausibly be an AI workload — interpreters, notebook kernels, and known agent binaries — matched by a case-insensitive candidate-runtime pattern. A compiler or a video encoder pegging a core is not an inference heartbeat, and treating it as one would make the signal worthless.
Local model runtimes¶
Twelve runtimes are recognised by process name and by the ports they own.
| Runtime | Ports | Exclusive |
|---|---|---|
| Ollama | 11434 | yes |
| LM Studio | 1234 | yes |
| GPT4All | 4891 | yes |
| Jan | 1337 | yes |
| llama.cpp server | 8080, 8081 | no |
| vLLM | 8000 | no |
| HF text-generation-inference | 8080 | no |
| Apple MLX LM server | 8080 | no |
| LocalAI | 8080 | no |
| Open WebUI | 3000, 8080 | no |
| KoboldCpp | 5001 | no |
| LiteLLM proxy | 4000 | no |
Withheld from the public documentation
Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.
Client detection¶
local_inference_client catches the more interesting half: not the model server, but whatever is calling it. A process holding a socket to 127.0.0.1:11434 is using a local model regardless of what it claims to be, and this is one of the few signals that survives a fully offline host with no egress at all.
Seven OpenAI-compatible request paths are recognised for shape matching:
/v1/chat/completions /v1/completions /v1/embeddings
/v1/responses /v1/messages
/api/generate /api/chat
Coverage without privilege¶
Unprivileged ps -Ao returns RSS and CPU time for the entire process table, across every user. So Plane A is essentially complete without root — which is why an unprivileged sensor can still tell you the machine is computing, even when it cannot tell you who it is talking to.
That asymmetry is the reason the startup banner reports coverage explicitly rather than leaving it to be inferred. See Startup checks.