← Playbook

Own your analytics, and the first thing you learn is how few of your visitors are people.

We stopped renting analytics and wrote a collector that runs on our own domain, enriches every event at the Cloudflare edge and stores raw rows in SQLite on a server we control. Having every row turned out to matter less for the dashboards than for the corrections: proxy networks posing as companies, mail scanners posing as readers, crawlers arriving by the back door. You can only fix what you can re-read.

2026-10-04/7 min read/Levelbrook AI Practice

The case for renting it

Hosted web analytics is one of the best deals in software. One script tag, no server, a polished dashboard, and for the most popular product no invoice at all. For a marketing site whose audience is the general public, most of what you need is there by lunchtime, and anyone who builds their own should be able to say why.

We can say why in three sentences. Our readers are engineers, hiring managers and technical buyers, the group most likely to run a content blocker, and a script that posts to a third-party analytics domain is the first thing a blocker removes. The hosted products hide individual rows behind thresholds and sampling, which is a reasonable privacy default and useless when the question is “did anyone from that company read the essay we sent them”. And when the product’s bot filtering is wrong, you cannot correct it, because you do not have the rows.

The third reason turned out to be the important one.

The path

Where each event goes. The only third party that sees it is the CDN that already serves the page.
Where each event goes. The only third party that sees it is the CDN that already serves the page.

The tracker is a 9 KB script with no dependencies. It records page views, engaged time, scroll depth, outbound clicks, downloads, form submissions (which field, never the value, and nothing that looks like a password or card number) and campaign tags from the URL. It posts to /e on the same domain as the page, so a blocker cannot drop it without breaking the site.

/e is a Cloudflare Pages Function, and this is the step that makes the whole design work. Cloudflare passes every Function a request.cf object, and in it is the name of the organisation that owns the visitor’s IP address, the ASN, the city, the TLS version, and whether Cloudflare has already classified the request as a known crawler. No script in the browser can learn any of that, and buying it as a GeoIP and ASN database is a recurring cost. The Function adds those fields, answers the browser immediately, and forwards the batch to our collector in the background, so a slow collector never slows a page.

The collector is a single Go binary using a pure-Go SQLite driver. It runs in a container capped at 192 MB and sits at about 12 MB resident on a server that hosts fifty other things. It classifies the network, flags bots, writes rows and serves a small dashboard. And because the store is one SQLite file, the most useful report is not the dashboard. It is a command that copies the entire database to a laptop over SSH and renders an offline page from it, with the raw .db alongside for any question the page does not answer.

What the rows taught us

375sessions in the first 13 days, across three sites
137of them not flagged as bots
~300fake company visits in 48 hours, from six proxy networks
1,364rows recorded under hostnames nobody links to

The first fortnight looked like a modest success: 375 sessions, 137 of them not obviously bots, a spike of 47 on the day an essay reached the Hacker News front page, and a Companies tab with recognisable-sounding organisations in it.

Then we read the rows. A large share of the “companies” were not companies. They were residential proxy exits, bulk hosting providers and scraping networks whose AS names do not contain the word hosting. In one 48-hour window, six such networks accounted for roughly three hundred sessions that had been bucketed as company visits, every one with zero engaged time and zero scroll. We added a proxy bucket to the classifier, and because the collector re-buckets stored rows at startup, the correction applied to all of history at once rather than from that day forward.

The second lesson was mail. Links in outbound email are followed, within seconds of delivery, by the recipient’s mail security product: Microsoft’s, Proofpoint’s, Mimecast’s, Barracuda’s. To a hosted analytics product those are visits, and to a hopeful sender they look like opens. They now have their own bucket, they are flagged, and they never count as engagement.

The third lesson was hostnames. Every Cloudflare Pages project answers on its own *.pages.dev name as well as its real domain, and crawlers find those names in certificate-transparency logs. 1,364 rows had arrived under hostnames that no human is ever sent to. The collector now folds them into the real site on ingest, backfills history at startup, and keeps the hostname the visitor actually used, because arriving by the back door is itself a strong crawler signal.

One number per visit, with its reasons

Inputs to the likely-human score in our offline report. Weights are ours and are shown in the report next to every visit.
Inputs to the likely-human score in our offline report. Weights are ours and are shown in the report next to every visit.

With every row available, each session gets a score from 0 to 100 and the reasons that produced it, visible on hover. The report defaults to showing likely humans only. The honest reading after that change was about four or five likely-human sessions a day across the estate at the time, which matched the hand audit that prompted it, and was a fraction of what any unfiltered count had said.

That is a humbling number, and it is the most valuable number the system has produced, because every decision that followed was made on it. The channel that reliably brought people (an essay in a community where engineers read) got more of our time. The channels that brought only scanners and crawlers got less. With a rented dashboard showing ten times the traffic, we would have spent that month optimising for paper silhouettes.

Answering the question that matters

The business question behind most of this is attribution. Did the person we wrote to read the thing we sent? Any link we send can carry a ?ref= tag; the tracker keeps it for the whole session, and the offline database answers the question with one query:

sql select date(ts/1000,'unixepoch') as day, as_org, path, engaged/1000 as seconds from events where ref_tag = 'acme-followup' and is_bot = 0 order by ts;

No sampling, no threshold, no export quota. If the answer is a mail scanner at delivery time and nothing after, that is an answer too.

What it costs to run, and the apps that are not on the edge

The running cost is close enough to zero to round. The collector shares a server with everything else and uses about 12 MB of it. The store is one SQLite file in a Docker volume, a few hundred megabytes after a month across the whole estate. The edge Functions run on Cloudflare’s free plan, whose daily request allowance (100,000 Function invocations a day at the time of writing) is far above what a set of sites like ours produces, and the tracker batches events so one page view is usually one request.

The interesting case is the applications that are not static sites: the Rails, Go and PHP apps on the servers. They cannot run a Pages Function of their own, and we did not want a collector endpoint on a separate subdomain that a blocker could learn. So they load the tracker from the main marketing domain and post to that domain’s /e, cross-origin, with a short allow-list of our own hostnames permitted. The browser still talks to Cloudflare directly, so the event still arrives with the network organisation attached, and the event carries its own hostname, so attribution stays with the app that was actually visited. One edge Function ends up serving about eighty hostnames across four hosting surfaces.

The privacy boundary

Owning the rows is also owning the responsibility for them. The system records what visitors do on pages we publish and which network they arrived from. It does not identify individuals, follow anyone across other sites, or record what anyone typed. For the product that holds private data, the tracker runs in a redact mode: no page titles, no link text, numeric ids and tokens in paths collapsed to :id and :token, nothing kept from query strings except campaign tags. The raw IP is stored next to a salted hash, and the retention decision is ours to make and to state.

And the data never leaves machines we control, apart from passing through the CDN that already serves the page. There is no processor to list in a privacy notice, no export to request, and nothing to migrate if a vendor changes its pricing or its product.

Where this is wrong

It costs time. The classifier is a set of regular expressions over AS organisation names, and it is only as good as the last audit. The dashboard is plain. Funnels, cohorts and the polished exploration a good analytics product offers are a SQL query away rather than a click away, which suits engineers and does not suit a marketing team. A site with a general audience and a team that does not read SQL should keep renting.

What to do on Monday

Pull a week of raw rows from whatever you use now, if you can, and look at the networks the sessions came from. If you cannot get the rows at all, that is the finding. If you can, sort by engaged time and count how many sessions have none. Then decide whether the number on your dashboard is the number you want to be making decisions on.

The tracker, the edge Function, the collector and the offline report are public as shoestring. It runs comfortably next to fifty other containers on a 2 GB server. The first thing it will show you is probably how few of your visitors are people, and that is worth knowing before you spend another month writing for the others.

Sources and things reacted to
  1. shoestring (GitHub, MIT): tracker, edge Function, collector and offline report the code described here, with our hostnames and credentials removed
  2. Cloudflare: the request.cf object asOrganization, asn, city, colo, tlsVersion and the verified-bot category used for enrichment
  3. modernc.org/sqlite the CGO-free SQLite driver that keeps the collector a single static binary
Keep reading

This is what we do all day.

Support automation, AI phone agents, n8n back-office work, and the engineering loop itself — always behind a gate you control.