Live Signals
Everything a search engine or an AI assistant does on this site is logged, and this page shows it: crawler visits, which engines arrive and for which pages, where real visitors come from, what the security layer blocks, and how all of it changes hour by hour and day by day.
We publish it for two reasons. Measurement is the part of search and AI visibility most agencies keep behind a login — we would rather show ours. And this is exactly the instrumentation we install on client sites, so what you see here is what you get.
What Each Card Shows
- Named crawler visits, last 24h (UTC) — automated requests the site received each hour. A bar appears only where traffic actually exists.
- All requests and top paths, last 24h (UTC) — every request in the window, including scanners, not-found requests and ordinary visitors, then the paths requested most. Named crawlers are the subset counted on card 1.
- Crawler mix today — which engines account for today’s crawler traffic, as a share of the total.
- Crawler share by engine — the same share measured across the window of logs available so far.
- Top engines in window — per-engine totals for that window, with each engine’s percentage.
- Window summary — engine, request count, distinct IPs and last-seen time for every crawler seen in that window.
- Latest crawler hits — the most recent requests: engine, path and time (UTC).
- Search & AI referrers (24h) — real visits that arrived from a search or AI source, identified from the HTTP referrer. “No referral traffic yet” simply means none has arrived.
- Last 8 hours by engine — hour-by-hour engine activity for the last eight hours.
- Security, last 24h — requests blocked by our crawler policy (HTTP 403) and scan attempts against paths such as /.env or /wp-admin.
- 404 ranking — the paths returning 404 most often; a useful technical SEO signal rather than a vanity metric.
- Top visitor IPs (masked) — the busiest sources with the last part of the address masked, plus the resolved country.
Crawler counts use requests from recognised search and AI engine user agents, and exclude requests to known scan paths (/.env, /.git, /wp-*, /index.php and similar). Raw counts including unrecognised agents are published in live.json with a matching definition field.
How This Data Is Collected
- Source: our own web server logs, with Cloudflare real-IP resolution so country data reflects the visitor rather than the CDN edge.
- Crawler identification: from each request’s user agent string. No analytics product, no tracking cookies and no client data are involved.
- Privacy: visitor IP addresses are never published in full — the last part is masked and counts are aggregated.
- Refresh: log parsing runs every five minutes; this page re-reads the data every sixty seconds.
- Retention: rotated logs limit the trend window while the site is young, so earlier days roll off as new ones arrive.
- Why the numbers are small: the site launched recently. New crawlers appear over time — we would rather show a real small number than an impressive invented one.
Want this instrumentation on your own site? It is how every engagement starts. Get in touch.
