Who's actually reading us
Parsed from Apache's access log: every search engine and AI crawler that has fetched a page in the last 7 days. The shape of this list is the entire SEO + GEO answer for an indie directory: are Googlebot, Bingbot, and the AI engines (ChatGPT, Perplexity, Claude) showing up? Bing-derived AI search surfaces fresh content fastest, which is why IndexNow and llms.txt are wired in. The chart below is the receipt.
Who's actually visiting
Real IP analysis: ip-api.com org lookup on non-bot IPs from Apache logs, then Haiku classifies the traffic composition. Updated weekly (Sundays).
The bot charts below cover declared crawlers (Googlebot says it's Googlebot). IP org lookup is the other half: what do the undeclared IPs actually resolve to?
From the last 7 days of Apache logs, roughly 38% of unique IPs resolve to cloud hosting ranges (OVH, AWS, DigitalOcean) — these are content scrapers pulling pages for training data or competitive indexing, not humans. About 8% hit residential proxy pools operated by rank-tracking services (Datacamp, Code200, M247): they send Google a search query to generate the GSC impression, then fetch the actual page through a proxy to confirm the ranking — which is why GSC impressions can spike on a page with barely any real clicks.
The 21% residential ISP traffic (Bell Canada / Sympatico leads the list) is the one category that could contain real people, though some of it is also the operator's own traffic before a DHCP lease change got filtered in GA4. The remaining 34% is unclassified — smaller regional ISPs, hosting blocks ip-api.com doesn't tag clearly. Weekly Haiku analysis runs Sunday mornings and writes the breakdown to a private cache file used by this section.
Top 10 crawlers, last 7 days
Every crawler, by type
Top 25 most-crawled pages, last 7 days
Search engines, share of crawl
AI & LLM crawlers, share of crawl
Bot farms, exploit-scanner probes by country
API spend
Classification, verification, and enrichment costs since —.
Daily spend
Snapshot: -