We stopped renting analytics and wrote a collector that runs on our own domain, enriches every event at the Cloudflare edge and stores raw rows in SQLite on a server we control. Having every row turned out to matter less for the dashboards than for the corrections: proxy networks posing as companies, mail scanners posing as readers, crawlers arriving by the back door. You can only fix what you can re-read.
Hosted web analytics is one of the best deals in software. One script tag, no server, a polished dashboard, and for the most popular product no invoice at all. For a marketing site whose audience is the general public, most of what you need is there by lunchtime, and anyone who builds their own should be able to say why.
We can say why in three sentences. Our readers are engineers, hiring managers and technical buyers, the group most likely to run a content blocker, and a script that posts to a third-party analytics domain is the first thing a blocker removes. The hosted products hide individual rows behind thresholds and sampling, which is a reasonable privacy default and useless when the question is “did anyone from that company read the essay we sent them”. And when the product’s bot filtering is wrong, you cannot correct it, because you do not have the rows.
The third reason turned out to be the important one.
The tracker is a 9 KB script with no dependencies. It records page views, engaged time, scroll depth,
outbound clicks, downloads, form submissions (which field, never the value, and nothing that looks
like a password or card number) and campaign tags from the URL. It posts to /e on the same domain as
the page, so a blocker cannot drop it without breaking the site.
/e is a Cloudflare Pages Function, and this is the step that makes the whole design work. Cloudflare
passes every Function a request.cf object, and in it is the name of the organisation that owns the
visitor’s IP address, the ASN, the city, the TLS version, and whether Cloudflare has already
classified the request as a known crawler. No script in the browser can learn any of that, and buying
it as a GeoIP and ASN database is a recurring cost. The Function adds those fields, answers the
browser immediately, and forwards the batch to our collector in the background, so a slow collector
never slows a page.
The collector is a single Go binary using a pure-Go SQLite driver. It runs in a container capped at
192 MB and sits at about 12 MB resident on a server that hosts fifty other things. It classifies the
network, flags bots, writes rows and serves a small dashboard. And because the store is one SQLite
file, the most useful report is not the dashboard. It is a command that copies the entire database to
a laptop over SSH and renders an offline page from it, with the raw .db alongside for any question
the page does not answer.
The first fortnight looked like a modest success: 375 sessions, 137 of them not obviously bots, a spike of 47 on the day an essay reached the Hacker News front page, and a Companies tab with recognisable-sounding organisations in it.
Then we read the rows. A large share of the “companies” were not companies. They were residential proxy exits, bulk hosting providers and scraping networks whose AS names do not contain the word hosting. In one 48-hour window, six such networks accounted for roughly three hundred sessions that had been bucketed as company visits, every one with zero engaged time and zero scroll. We added a proxy bucket to the classifier, and because the collector re-buckets stored rows at startup, the correction applied to all of history at once rather than from that day forward.
The second lesson was mail. Links in outbound email are followed, within seconds of delivery, by the recipient’s mail security product: Microsoft’s, Proofpoint’s, Mimecast’s, Barracuda’s. To a hosted analytics product those are visits, and to a hopeful sender they look like opens. They now have their own bucket, they are flagged, and they never count as engagement.
The third lesson was hostnames. Every Cloudflare Pages project answers on its own *.pages.dev name
as well as its real domain, and crawlers find those names in certificate-transparency logs. 1,364 rows
had arrived under hostnames that no human is ever sent to. The collector now folds them into the real
site on ingest, backfills history at startup, and keeps the hostname the visitor actually used,
because arriving by the back door is itself a strong crawler signal.
With every row available, each session gets a score from 0 to 100 and the reasons that produced it, visible on hover. The report defaults to showing likely humans only. The honest reading after that change was about four or five likely-human sessions a day across the estate at the time, which matched the hand audit that prompted it, and was a fraction of what any unfiltered count had said.
That is a humbling number, and it is the most valuable number the system has produced, because every decision that followed was made on it. The channel that reliably brought people (an essay in a community where engineers read) got more of our time. The channels that brought only scanners and crawlers got less. With a rented dashboard showing ten times the traffic, we would have spent that month optimising for paper silhouettes.
The business question behind most of this is attribution. Did the person we wrote to read the thing
we sent? Any link we send can carry a ?ref= tag; the tracker keeps it for the whole session, and the
offline database answers the question with one query:
sql
select date(ts/1000,'unixepoch') as day, as_org, path, engaged/1000 as seconds
from events
where ref_tag = 'acme-followup' and is_bot = 0
order by ts;
No sampling, no threshold, no export quota. If the answer is a mail scanner at delivery time and nothing after, that is an answer too.
The running cost is close enough to zero to round. The collector shares a server with everything else and uses about 12 MB of it. The store is one SQLite file in a Docker volume, a few hundred megabytes after a month across the whole estate. The edge Functions run on Cloudflare’s free plan, whose daily request allowance (100,000 Function invocations a day at the time of writing) is far above what a set of sites like ours produces, and the tracker batches events so one page view is usually one request.
The interesting case is the applications that are not static sites: the Rails, Go and PHP apps on the
servers. They cannot run a Pages Function of their own, and we did not want a collector endpoint on a
separate subdomain that a blocker could learn. So they load the tracker from the main marketing domain
and post to that domain’s /e, cross-origin, with a short allow-list of our own hostnames permitted.
The browser still talks to Cloudflare directly, so the event still arrives with the network
organisation attached, and the event carries its own hostname, so attribution stays with the app that
was actually visited. One edge Function ends up serving about eighty hostnames across four hosting
surfaces.
Owning the rows is also owning the responsibility for them. The system records what visitors do on
pages we publish and which network they arrived from. It does not identify individuals, follow anyone
across other sites, or record what anyone typed. For the product that holds private data, the tracker
runs in a redact mode: no page titles, no link text, numeric ids and tokens in paths collapsed to
:id and :token, nothing kept from query strings except campaign tags. The raw IP is stored next to
a salted hash, and the retention decision is ours to make and to state.
And the data never leaves machines we control, apart from passing through the CDN that already serves the page. There is no processor to list in a privacy notice, no export to request, and nothing to migrate if a vendor changes its pricing or its product.
It costs time. The classifier is a set of regular expressions over AS organisation names, and it is only as good as the last audit. The dashboard is plain. Funnels, cohorts and the polished exploration a good analytics product offers are a SQL query away rather than a click away, which suits engineers and does not suit a marketing team. A site with a general audience and a team that does not read SQL should keep renting.
Pull a week of raw rows from whatever you use now, if you can, and look at the networks the sessions came from. If you cannot get the rows at all, that is the finding. If you can, sort by engaged time and count how many sessions have none. Then decide whether the number on your dashboard is the number you want to be making decisions on.
The tracker, the edge Function, the collector and the offline report are public as shoestring. It runs comfortably next to fifty other containers on a 2 GB server. The first thing it will show you is probably how few of your visitors are people, and that is worth knowing before you spend another month writing for the others.
Support automation, AI phone agents, n8n back-office work, and the engineering loop itself — always behind a gate you control.