Browser telemetry¶
Not implemented
There is no browser module today. This page explains a real and known gap in coverage so you can reason about it, and sketches the direction. Nothing described here as "proposed" exists in the shipped product.
The gap¶
A great deal of shadow AI never touches a process the sensor can usefully name. Someone opens chatgpt.com in a tab, pastes a customer list into it, and reads the answer. From the endpoint, what happened was:
The socket is real, the peer is catalogued, and a finding is raised — shadow_ai_egress at the consumer-chat weight of 30. What the sensor cannot see is anything that matters after that:
| Question | Answerable from the endpoint? |
|---|---|
| Which browser process opened it? | Yes. |
| Which provider? | Yes, if the host is catalogued or inference-shaped. |
| Which tab, profile, or extension? | No. |
| Which user account at that provider? | No. |
| What was sent? | No. |
| Was a file uploaded? | No. |
Why this is structural, not a bug¶
TLS terminates inside the browser. Below that boundary there is a socket to a CDN address and nothing else. No amount of improvement to lsof parsing, DNS capture, or Endpoint Security subscription changes that, because the information is not present at any of those layers.
Compounding it:
- Process attribution is coarse. Chrome's multi-process model means the socket often belongs to a helper or network-service process, not to anything a human would call "the browser tab".
- Shared CDN addresses. Consumer AI properties sit on the same Cloudflare and Google ranges as everything else those companies serve, so catalog-based attribution is weaker here than for a dedicated API host.
- DNS-over-HTTPS is on by default in several browsers, which defeats the DNS sniffer for exactly the traffic this concerns.
The behavioural planes help less than usual, too. A browser always burns CPU and always holds many TLS connections open, so Plane A signals that discriminate well for an interpreter discriminate poorly for Chrome.
What that means today¶
Treat browser-mediated AI use as detected but not characterised.
You will see that a host reached a consumer AI property, when, and from which browser. You will not see who, in which account, or with what data. For an inventory question — is anyone using AI here? — that is often enough. For an incident question — what left the building? — it is not, and no configuration change closes the gap.
The four consumer chat providers in the catalog are weighted at 30 rather than the 50 a frontier API carries, partly for this reason: the observation is real but thin.
Direction, if it is ever built¶
Two capabilities that sound like one thing and are not, and separating them is the main design conclusion.
Capability 1
Metadata
Which AI origin, which browser profile, which extension, how often. No message content.
- Better Plane B attribution
- Modest privacy exposure
- The plausible first target
Capability 2
Content
What was typed, pasted, or uploaded.
- Answers the incident question
- Enormous privacy and consent exposure
- A separate product decision
Boundary
New attack surface
A native messaging host would be the sensor's first untrusted-input boundary.
- Everything today is read-only acquisition
- An extension channel is bidirectional
- Needs its own threat model
A plausible sequencing would be a Manifest V3 metadata extension for Chrome, Edge and Brave first; Firefox as a delta; Safari sharing the signed-app vehicle with any future Network Extension. Content capture would be a distinct, separately-consented capability, not a flag on the same extension.
What should come first¶
Browser work is not the highest-value next thing, and saying so is part of the analysis. Several cheaper fixes improve real coverage today:
- UDP and QUIC acquisition. A growing share of AI traffic is HTTP/3, and the current socket acquisition is TCP-shaped.
- Browser process classification. Recognising helper and network-service processes as browser rather than as generic candidates would stop them competing with genuine interpreter signals.
- Consumer-chat weighting. Revisit the 30 now that the attribution quality behind it is written down.
- Catalog coverage. More consumer properties, and better handling of the shared-CDN false-attribution case documented in Known limitations.