# What actually reads this site Most of this site's traffic is machines. I show which ones, what they ask for, and how easily a chart about them can be poisoned, with every figure's window and method beside it. > Source: https://eduarddziak.com/visibility/ > Window: 2026-08-13 to 2026-08-23, 11 days > Updated: 2026-09-01 > Data: /data/visibility.json and /data/visibility.csv > Licence: CC BY 4.0 The primary window excludes 22 August 2026, the day one hour of requests claimed crawler names that belong to 5 operators. Every figure names its own window, and Cloudflare-derived figures are estimates under sampling. ## Summary | Measure | Value | Window | | --- | --- | --- | | AI share of page views, with 2026-08-22 | 16.4% | 11 days to 23 Aug 2026 | | AI share of page views, without 2026-08-22 | 3.5% | 10 days to 23 Aug 2026 | | 404 share of page requests, without 2026-08-22 | 50.5% | 10 days to 23 Aug 2026 | | Named-query impressions | 0.8% | 21 days to 18 Aug 2026 | One hour multiplied the apparent AI share 4.7 times. ## Key findings - One hour of requests claiming 7 AI crawler names lifted this site's apparent AI crawler share from 3.5% to 16.4%. - On 22 August 2026, page views from clients claiming AI crawler names were 49.0% of that day's page views. The next day they were 0.5%. - 0 of 140 requests claiming ChatGPT-User on 22 August 2026 came from OpenAI's published addresses, and all 140 came from a single address. - Googlebot verified 40 of 40 from 8 addresses on 22 August 2026. - Excluding 22 August 2026, 50.5% of page requests asked for pages this site never served. - 84.3% of unrecognised page views claim to be Chrome, and 43.9% of those claim majors released before 2024. - Google named the search query behind 2 of 245 impressions over the 21 days it had finished revising, which is 0.8%. ## Without 22 August 2026, 77.5% of page views come from clients I cannot name | Bucket | Share without 2026-08-22 | Share with 2026-08-22 | | --- | --- | --- | | AI crawlers | 3.5% | 16.4% | | Search engine crawlers | 7.6% | 6.1% | | SEO tool crawlers | 9.4% | 6.9% | | Other recognised bots and tools | 2.1% | 1.6% | | Unrecognised clients, not verifiable as people | 77.5% | 68.9% | One bad hour moves the AI row from 3.5% to 16.4%. I expected the wave to show in the AI row. I did not expect the row to move this far. If you publish an AI-traffic share from user-agent strings alone, one hour like this can move yours just as far. Nobody reading your chart would know. The last row is what my classification list did not recognise. I do not present it as people, and the section "What the unrecognised traffic claims to be" says why. ## The hour that poisoned the chart On 22 August 2026, 557 requests arrived claiming 7 crawler names that belong to 5 operators, OpenAI, Perplexity, Google, Anthropic and Amazon. 539 of them landed in a single hour. All of them came from one country. 95.0% were answered 404, meaning they asked for pages that do not exist here. The obvious question is who sent them. On the evidence I keep, spoofing is the near-certain reading and cannot be proven, because proving it needs the IP addresses behind those requests, and I refuse to hold them. No operator named here is accused of anything. The names were the costume, and an operator whose name was worn is the party imitated, not the party acting. If the same hour had landed in your logs, your AI share for the whole month would carry it. ### On 22 August 2026, 49.0% of page views came from clients claiming an AI crawler's name. The next day, 0.5% Share of that day's page views by claimed client, 11 days, 13 Aug 2026 to 23 Aug 2026. AI means AI crawlers, Search means Search engine crawlers, SEO means SEO tool crawlers, and Other means Other recognised bots and tools. | Day | AI | Search | SEO | Other | Unrecognised clients, not verifiable as people | Note | | --- | --- | --- | --- | --- | --- | --- | | 13 Aug 2026 | 0.9% | 0.0% | 3.6% | 1.8% | 93.7% | Site release day. | | 14 Aug 2026 | 0.0% | 2.7% | 6.6% | 0.5% | 90.2% | | | 15 Aug 2026 | 0.8% | 3.1% | 3.9% | 0.8% | 91.4% | | | 16 Aug 2026 | 0.0% | 5.2% | 5.2% | 0.7% | 88.9% | | | 17 Aug 2026 | 12.4% | 3.3% | 6.6% | 0.0% | 77.7% | | | 18 Aug 2026 | 4.8% | 16.3% | 33.7% | 0.7% | 44.6% | Site release day. | | 19 Aug 2026 | 2.5% | 3.8% | 1.9% | 2.5% | 89.4% | Site release day. | | 20 Aug 2026 | 3.1% | 17.1% | 3.1% | 8.8% | 68.0% | | | 21 Aug 2026 | 13.8% | 16.8% | 23.9% | 4.6% | 41.0% | | | 22 Aug 2026 | 49.0% | 2.3% | 0.7% | 0.6% | 47.5% | The wave day. Seven crawler names in one hour. | | 23 Aug 2026 | 0.5% | 4.3% | 1.1% | 0.8% | 93.4% | Second burst day. Requests for pages that never existed. | Daily shares on a site this small move a lot day to day, and the figures are estimates under sampling. I keep coming back to that pair. One day that high and the next that low is not a trend. It is one event. A monthly average would keep the lift and hide that it came from one day. That shape is the first thing I would look for in your own daily numbers before trusting any monthly AI figure. ### One hour multiplied the apparent AI share 4.7 times | Measure | Value | | --- | --- | | Crawler names claimed | ChatGPT-User, GPTBot, OAI-SearchBot, PerplexityBot, Google-Extended, ClaudeBot, Amazonbot | | Operators whose names were used | 5 | | Requests across those names | 557 | | Landed in the single peak hour | 539 | | Answered 404 | 95.0% | | Countries of origin | 1 | Why publish the AI share twice? Because one hour moved it. Any chart that trusts user-agent strings can be moved the same way, in one hour. If you quote only the higher figure from your own logs, you are quoting the hour, not the site. ### Checked against the operators' published ranges, 22 August 2026 Claiming ChatGPT-User, 0 of 140 requests came from OpenAI's published ranges, and all 140 came from a single address. Googlebot, the same day, 40 of 40 came from Google's published ranges, from 8 addresses. This chart covers only the crawlers whose operators publish address ranges. The day's largest claimant, Amazonbot at 185 requests, has no published ranges to check against, so it appears in the table with the verdict Unverifiable. | Crawler claimed | Operator | Requests | Verified | Distinct addresses | Verdict | | --- | --- | --- | --- | --- | --- | | Googlebot | Google | 40 | 40 | 8 | Verified | | bingbot | Microsoft | 8 | 8 | 8 | Verified | | GoogleOther | Google | 4 | 4 | 2 | Verified | | OAI-SearchBot | OpenAI | 48 | 3 | 4 | 3 of 48 verified | | GPTBot | OpenAI | 49 | 1 | 2 | 1 of 49 verified | | ChatGPT-User | OpenAI | 140 | 0 | 1 | Not verified | | PerplexityBot | Perplexity | 46 | 0 | 1 | Not verified | | Amazonbot | Amazon | 185 | Unverifiable | Unverifiable | Unverifiable. This operator publishes no ranges, so no check is possible. | | Google-Extended | Google | 46 | Not a user agent | 1 | Not a user agent. Google documents Google-Extended as a robots.txt control token with no HTTP user agent string, so a request carrying it in its user agent did not come from Google. ([Google's crawler documentation](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers)) | | ClaudeBot | Anthropic | 43 | Unverifiable | Unverifiable | Unverifiable. This operator publishes no ranges, so no check is possible. | Counts here are requests, not page views. Verification filters to eyeball traffic but applies neither the page-view rule nor the own-tooling exclusion. Per the review, verdicts publish from this day only. Other covered days carry partial failures whose addresses are discarded by design and can never be re-examined. Two different claims sit in that table, and I keep them apart. That the failing requests came from outside the ranges their operators publish is measured, and the check ran against a snapshot of the ranges taken that same day. That somebody wore those names as a costume is a reading, and it stays a reading whatever later data shows. Was the check itself working? Yes, and the proof is in the same table. OpenAI's own range file verified other requests on the day it verified 0 of 140 for ChatGPT-User. A range file that verifies nothing cannot be told apart from a broken lookup. One that verifies some requests and not others can. That control is the part I would build first if this were your site, because without it a zero is a number anyone can wave away. ## 84.3% of the unrecognised traffic claims to be Chrome, much of it years old I call this bucket unrecognised because my list did not match it. That is all I know. I cannot tell you these are people, and I will not pretend to. 84.3% of unrecognised page views claim to be Chrome. Of those Chrome claims, 43.9% name major versions released before 2024. Chrome 120 went stable on 2023-12-05 and the next major not until 2024-01-23, so a claim of major 120 or older is a claim of a pre-2024 browser. A client claiming Chrome 78, released 2019-10-22, appeared on 11 of 11 observed days. A real browser years out of date is possible. That much of a site's unknown traffic running one is not what I would expect from people. A Chrome 78 client that shows up on all the observed days settles it for me. If your analytics counts these as visitors, that is the number I would doubt first. | Claimed major | Stable release date | Date source | Share of Chrome-claiming page views | Note | | --- | --- | --- | --- | --- | | Chrome 131 | 2024-11-12 | [Chromium Dash release schedule](https://chromiumdash.appspot.com/fetch_milestone_schedule?mstone=131) | 39.8% | Mostly the 2026-08-22 and 2026-08-23 bursts | | Chrome 78 | 2019-10-22 | [Chromium Dash release schedule](https://chromiumdash.appspot.com/fetch_milestone_schedule?mstone=78) | 28.0% | Present on 11 of 11 observed days | | Chrome 89 | 2021-03-02 | [Chromium Dash release schedule](https://chromiumdash.appspot.com/fetch_milestone_schedule?mstone=89) | 8.9% | | | Chrome 151 | 2026-07-28 | [Chromium Dash release schedule](https://chromiumdash.appspot.com/fetch_milestone_schedule?mstone=151) | 4.9% | | ## Most page requests ask for a page this site never served Each window's figure is the share of page requests answered 404, and a 404 counts because this site answers it with a real HTML page. Half surprised me. I expected requests for pages that do not exist to be a background hum, not the loudest thing in the log. If your site is small, I would expect yours to look similar. A 404 share this high says more about the requests arriving than about your own broken links. Most of the 404 numerator comes from unrecognised clients, so this is a claim about request behaviour, never about who is behind it. | Window | Share answered 404 | | --- | --- | | without 2026-08-22 | 50.5% | | all eleven days | 62.5% | | without 2026-08-22 and 2026-08-23 | 41.0% | ## Who actually reads the markdown mirror Every page here has a markdown twin at the same address ending in .md, and each row is one bucket's share of the fetches of those twins. For ClaudeBot, the markdown mirror made up 45.7% of its own content fetches in this window. That one I did not see coming. I built the mirror for readers like it and still expected the HTML to win by far more than it did. If you publish a plain-text twin of your pages, the same split in your own logs is what would tell you whether it earned its keep. Shares describe observed readers, not appetite, because a client that never found the mirror could not have read it, and the AI slice rests on few enough fetches that I hold it at first-signals strength. | Bucket | Share of markdown fetches | | --- | --- | | SEO tool crawlers | 50.0% | | Unrecognised clients, not verifiable as people | 18.1% | | Search engine crawlers | 16.2% | | AI crawlers | 15.3% | | Other recognised bots | 0.5% | ## Who checks robots.txt before taking content? Both columns divide a crawler's robots.txt fetches in this window, first by the HTML pages it took, then by its content fetches with the markdown mirror counted as content. Applebot, ClaudeBot and OAI-SearchBot fetched robots.txt more often than they took a page, and for 7 of the 13 names no robots.txt fetch was observed in 10 days. The zeros are the rows I read most carefully. No robots.txt fetch observed in 10 days is not proof of no fetch, because a crawler that read the file once, before the window opened, would look the same here. The ratios are the part you can use, since they say which names check often before taking content from a site like yours. Every name is self-identified, and etiquette observed is not obedience proven. | Crawler | Robots.txt fetches per HTML page taken (10 days, without 2026-08-22) | Per content fetch, markdown counted as content | | --- | --- | --- | | Applebot | 4.0 | 0.8 | | ClaudeBot | 3.7 | 2.0 | | OAI-SearchBot | 1.5 | 1.4 | | Googlebot | 0.7 | 0.6 | | bingbot | 0.1 | 0.1 | | PerplexityBot | 0.1 | 0.1 | | GoogleOther | No robots.txt fetch observed in 10 days. | No robots.txt fetch observed in 10 days. | | GPTBot | No robots.txt fetch observed in 10 days. | No robots.txt fetch observed in 10 days. | | ChatGPT-User | No robots.txt fetch observed in 10 days. | No robots.txt fetch observed in 10 days. | | Amazonbot | No robots.txt fetch observed in 10 days. | No robots.txt fetch observed in 10 days. | | YandexBot | No robots.txt fetch observed in 10 days. | No robots.txt fetch observed in 10 days. | | CCBot | No robots.txt fetch observed in 10 days. | No robots.txt fetch observed in 10 days. | | Baiduspider | No robots.txt fetch observed in 10 days. | No robots.txt fetch observed in 10 days. | ## Google named the search query behind 2 of 245 impressions Across 21 settled days, Google named the search query behind 2 of 245 impressions, which is 0.8%. Query rows are a floor, never a total, because the gap is queries Google declined to name. The withholding is a known mechanism, not something odd about this site. [Ahrefs measured it at 46.77%](https://ahrefs.com/blog/gsc-anonymized-queries/) across the sites it studied, so a site this small sits at the far end of the same mechanism, where almost all of the queries disappear. What I watch is how the share moves, so the table gains a row at each refresh. If your site is small too, the movement in your named share is the part worth keeping, not the level. | Window end | Settled days | Impressions with a named query | All impressions | Share | | --- | --- | --- | --- | --- | | 2026-08-18 | 21 | 2 | 245 | 0.8% | ## First signals, and what I am holding back First signals are observations too few to claim a pattern, published with their size stated so nobody mistakes them for one. Each finding below waits on the condition beside it, and leaves this list only through a review round. | Register id | Held finding | Publishes when | | --- | --- | --- | | VIS-010 | llms.txt before and after | Three to four weeks of after-data; the before-baseline and go-live date are already recorded. | | VIS-008 | Crawl versus credit, Google edition | All joint Search Console days settled. | | VIS-009 | AI training / search / user-fetch split | AI page views in the hundreds; one crawler session currently moves the split by 24 points. | ## How this is measured Two archives sit behind this page, and I ask them different questions. The traffic archive holds sampled Cloudflare analytics for this site and tells me what asked for what. The verification archive holds daily counts from checking crawler-name claims against published address ranges and tells me whether a name's addresses matched what its operator publishes. Google Search Console is the third source. It tells me what Google's index did with the site afterwards. Four rules shape the traffic figures, and you need them to check my work. I count eyeball traffic only, which is Cloudflare's own term for requests arriving from outside its network. A page view is a request for a page-shaped path, meaning any address that is not an asset such as a stylesheet, font or image, answered 2xx or 404, because a 404 on this site returns a real HTML page, and counting successes only would miss most of what actually happens here. I set that rule by checking it against Cloudflare's own page-view total for the same days. My own fetches and this project's tooling come out before anything is counted. Finally, Cloudflare samples this data, so the figures derived from it are estimates, with sampling intervals up to 3.3 observed, and none of them is an exact tally. If you recompute a figure and land a little off mine, sampling is the first place I would look. Everything my classification list does not match lands in one bucket, labelled "Unrecognised clients, not verifiable as people". I do not present that bucket as human traffic in any chart, table or download. The reason is in the section on what that traffic claims to be. A client name on this page is the name the client gave itself. A name is not proof of who sent the request, and no operator is accused of anything here. Since the verification collector first ran on 24 August 2026, I compare any request claiming a known crawler's name, as it is collected, against the address ranges that crawler's operator publishes. The comparison happens in memory. Only counts are stored, and the addresses are discarded. I snapshot the ranges the same day, so when you read a verdict it is against exactly what the operator published then, not what it publishes now. Five verdict words appear on this page, and they are not interchangeable. Verified means the requests came from inside the operator's published ranges. Not verified means the operator publishes ranges and the requests came from outside them. Unverifiable means the operator publishes no ranges at all, so no check is possible. That is a fact about the operator's documentation, never a verdict on the crawler, and it is why Amazon's Amazonbot and Anthropic's ClaudeBot can show no number. Not covered means the day sits outside verification's reach. Not a user agent means the name is documented as never being an HTTP user agent at all, so no address range could redeem it. Verification cannot prove intent. It shows that an address sat inside or outside a published range, and that part is measured. Calling the rest a costume carries the same limit as the verdict panel above, and I keep the two apart everywhere on this page. The check also reaches back only seven days from its first run, so days before that reach are unverified permanently. No refresh may add a verdict to an earlier day, and no refresh may turn the reading into a proven claim, whatever later days show. Some things I deliberately do not collect or publish anywhere on this site. IP addresses. Query strings. Referrers. Bot scores. Raw request paths. Raw user-agent strings. A count of distinct addresses is a count, and it is the only address-shaped thing I will ever publish here. You will not find visitor counts here. My site is small, and a raw count would make you judge the size instead of the pattern, which is the part you can use. An event count appears only as method context inside one named episode, with its denominator and its window in the same block. Search Console figures count settled days only. Google keeps revising its most recent days, so I leave an unsettled day out entirely instead of showing it provisionally. Days that look odd have public explanations. 22 August 2026 carries the wave hour. 23 August 2026 carried a second burst of requests for pages that do not exist, which is why the 404 table names a window without both days. And on 24 August 2026 I changed the site's robots policy, added a tripwire behind it with a block for Bytespider, and took /llms.txt live, so the launch window ends the day before and the etiquette table is a before-baseline. I refresh these figures through the same review round that produced them, and each refresh adds a row to the search series and the after-policy columns to the robots table. 11 days of detailed data, 13 August 2026 to 23 August 2026. 21 settled Search Console days. Built 1 September 2026. The figures come from this site's own Cloudflare zone analytics and from Google Search Console, read through collectors I run myself. Chrome release dates come from the Chromium Dash schedule service.