Visibility
What actually reads this site
Most of this site's traffic is machines. I can show you which ones, what they ask for, and how easily a chart about them can be poisoned. Where an operator publishes the addresses its crawler uses, I check the claim against them, so some of what follows is more than the name a client gave itself. The figures below carry their own observation windows and their method beside them. You can check any of them against the downloads at the end, and I can recompute any one of them from the archive it came from.
One day is treated specially throughout. On 22 August 2026, one hour of requests claimed crawler names that belong to 5 operators, so the primary window excludes that day, and the strips below say either all 11 days or without 2026-08-22. The hour that poisoned the chart tells the full story.
- AI share, with 2026-08-2216.4%
- AI share, without 2026-08-223.5%
- 404 share, without 2026-08-2250.5%
- Named-query impressions0.8%
The dashed cell is the reading a naive chart would have published.
Key findings
- One hour of requests claiming 7 AI crawler names lifted this site's apparent AI crawler share from 3.5% to 16.4%.
- On 22 August 2026, page views from clients claiming AI crawler names were 49.0% of that day's page views. The next day they were 0.5%.
- 0 of 140 requests claiming ChatGPT-User on 22 August 2026 came from OpenAI's published addresses, and all 140 came from a single address.
- Googlebot verified 40 of 40 from 8 addresses on 22 August 2026.
- Excluding 22 August 2026, 50.5% of page requests asked for pages this site never served.
- 84.3% of unrecognised page views claim to be Chrome, and 43.9% of those claim majors released before 2024.
- Google named the search query behind 2 of 245 impressions over the 21 days it had finished revising, which is 0.8%.
Figures cover 11 days to 23 Aug 2026 unless a line names its own window. Cloudflare-derived figures are estimates under sampling.
Without 22 August 2026, 77.5% of page views come from clients I cannot name
11 days of detailed data · to 23 Aug 2026 · sampled, estimates not talliesShare of page views. A page view is a request for a page address answered 2xx or 404, from clients outside Cloudflare's own network, with my own tooling removed first.
Each row splits page views by what the client claimed to be, first without 22 August 2026, then with it, dashed.
One bad hour moves the AI row from 3.5% to 16.4%. I expected the wave to show in the AI row. I did not expect the row to move this far. If you publish an AI-traffic share from user-agent strings alone, one hour like this can move yours just as far. Nobody reading your chart would know.
| Bucket | Share without 2026-08-22 | Share with 2026-08-22 |
|---|---|---|
| AI crawlers | 3.5% | 16.4% |
| Search engine crawlers | 7.6% | 6.1% |
| SEO tool crawlers | 9.4% | 6.9% |
| Other recognised bots and tools | 2.1% | 1.6% |
| Unrecognised clients, not verifiable as people why this is not called users | 77.5% | 68.9% |
The same figures are published as JSON and CSV, so you can check any number on this page against them.
The hour that poisoned the chart
11 days of detailed data · to 23 Aug 2026 · sampled, estimates not talliesShare of page views. A page view is a request for a page address answered 2xx or 404, from clients outside Cloudflare's own network, with my own tooling removed first.
On 22 August 2026, 557 requests arrived claiming 7 crawler names that belong to 5 operators, OpenAI, Perplexity, Google, Anthropic and Amazon. 539 of them landed in a single hour. All of them came from one country. 95.0% were answered 404, meaning they asked for pages that do not exist here.
The obvious question is who sent them. On the evidence I keep, spoofing is the near-certain reading and cannot be proven, because proving it needs the IP addresses behind those requests, and I refuse to hold them. No operator named here is accused of anything. The names were the costume, and an operator whose name was worn is the party imitated, not the party acting. If the same hour had landed in your logs, your AI share for the whole month would carry it.
On 22 August 2026, 49.0% of page views came from clients claiming an AI crawler's name. The next day, 0.5%
Each point is the share of that day's page views that came from clients claiming an AI crawler's name.
The line sits in single digits on most days, and on 22 Aug 2026 it reaches 49.0% of that day's page views. One day later it is 0.5%. I keep coming back to that pair. One day that high and the next that low is not a trend. It is one event. A monthly average would keep the lift and hide that it came from one day. That shape is the first thing I would look for in your own daily numbers before trusting any monthly AI figure.
| Day | AI | Search | SEO | Other | Unrecognised clients, not verifiable as people | Note |
|---|---|---|---|---|---|---|
| 13 Aug 2026 | 0.9% | 0.0% | 3.6% | 1.8% | 93.7% | Site release day. |
| 14 Aug 2026 | 0.0% | 2.7% | 6.6% | 0.5% | 90.2% | |
| 15 Aug 2026 | 0.8% | 3.1% | 3.9% | 0.8% | 91.4% | |
| 16 Aug 2026 | 0.0% | 5.2% | 5.2% | 0.7% | 88.9% | |
| 17 Aug 2026 | 12.4% | 3.3% | 6.6% | 0.0% | 77.7% | |
| 18 Aug 2026 | 4.8% | 16.3% | 33.7% | 0.7% | 44.6% | Site release day. |
| 19 Aug 2026 | 2.5% | 3.8% | 1.9% | 2.5% | 89.4% | Site release day. |
| 20 Aug 2026 | 3.1% | 17.1% | 3.1% | 8.8% | 68.0% | |
| 21 Aug 2026 | 13.8% | 16.8% | 23.9% | 4.6% | 41.0% | |
| 22 Aug 2026 | 49.0% | 2.3% | 0.7% | 0.6% | 47.5% | The wave day. Seven crawler names in one hour. |
| 23 Aug 2026 | 0.5% | 4.3% | 1.1% | 0.8% | 93.4% | Second burst day. Requests for pages that never existed. |
Daily shares on a site this small move a lot day to day, and the figures are estimates under sampling.
One hour multiplied the apparent AI share 4.7 times
Why publish the AI share twice? Because one hour moved it. Any chart that trusts user-agent strings can be moved the same way, in one hour. If you quote only the higher figure from your own logs, you are quoting the hour, not the site.
| Measure | Value |
|---|---|
| Crawler names claimed | ChatGPT-User, GPTBot, OAI-SearchBot, PerplexityBot, Google-Extended, ClaudeBot, Amazonbot |
| Operators whose names were used | 5 |
| Requests across those names | 557 requests |
| Landed in the single peak hour | 539 requests |
| Answered 404 | 95.0% |
| Countries of origin | 1 |
Checked against the operators' published ranges, 22 August 2026
2026-08-22 · verification archive · requests, not page viewsCounts here are requests, not page views. Verification filters to eyeball traffic but applies neither the page-view rule nor the own-tooling exclusion.
This chart covers only the crawlers whose operators publish address ranges. The day's largest claimant, Amazonbot at 185 requests, has no published ranges to check against, so it appears in the table with the verdict Unverifiable.
| Crawler claimed | Operator | Requests | Verified | Distinct addresses | Verdict |
|---|---|---|---|---|---|
| Googlebot | 40 | 40 | 8 | Verified | |
| bingbot | Microsoft | 8 | 8 | 8 | Verified |
| GoogleOther | 4 | 4 | 2 | Verified | |
| OAI-SearchBot | OpenAI | 48 | 3 | 4 | 3 of 48 verified |
| GPTBot | OpenAI | 49 | 1 | 2 | 1 of 49 verified |
| ChatGPT-User | OpenAI | 140 | 0 | 1 | Not verified |
| PerplexityBot | Perplexity | 46 | 0 | 1 | Not verified |
| Amazonbot | Amazon | 185 | Unverifiable. This operator publishes no ranges, so no check is possible. | ||
| Google-Extended | 46 | Not a user agent | 1 | Not a user agent. Google documents Google-Extended as a robots.txt control token with no HTTP user agent string, so a request carrying it in its user agent did not come from Google. Google's crawler documentation | |
| ClaudeBot | Anthropic | 43 | Unverifiable. This operator publishes no ranges, so no check is possible. | ||
Two different claims sit in that table, and I keep them apart. That the failing requests came from outside the ranges their operators publish is measured, and the check ran against a snapshot of the ranges taken that same day. That somebody wore those names as a costume is a reading, and it stays a reading whatever later data shows. Was the check itself working? Yes, and the proof is in the same table. OpenAI's own range file verified other requests on the day it verified 0 of 140 for ChatGPT-User. A range file that verifies nothing cannot be told apart from a broken lookup. One that verifies some requests and not others can. That control is the part I would build first if this were your site, because without it a zero is a number anyone can wave away.
Per the review, verdicts publish from this day only. Other covered days carry partial failures whose addresses are discarded by design and can never be re-examined.
Verification first ran on 24 August 2026 and could reach back only seven days, so earlier days, including the start of the traffic window on 13 August 2026, are unverified and will stay that way. No later refresh can change that.
84.3% of the unrecognised traffic claims to be Chrome, much of it years old
11 days of detailed data · to 23 Aug 2026 · sampled, estimates not talliesClaimed Chrome major versions among unrecognised page views. Each bar's denominator is the Chrome-claiming slice, named beside the chart.
I call this bucket unrecognised because my list did not match it. That is all I know. I cannot tell you these are people, and I will not pretend to.
One bar per claimed Chrome major version, as a share of the unrecognised page views that claim Chrome.
A client claiming Chrome 78, released 22 October 2019, appeared on 11 of 11 observed days.
84.3% of unrecognised page views claim to be Chrome. Of those Chrome claims, 43.9% name major versions released before 2024. Chrome 120 went stable on 5 December 2023 and the next major not until 23 January 2024, so a claim of major 120 or older is a claim of a pre-2024 browser. A real browser years out of date is possible. That much of a site's unknown traffic running one is not what I would expect from people. A Chrome 78 client that shows up on all the observed days settles it for me. If your analytics counts these as visitors, that is the number I would doubt first.
The Chrome 131 bar is dashed because its spike is largely the 22 August 2026 and 23 August 2026 bursts.
| Claimed major | Stable release | Date source | Share of Chrome-claiming page views | Note |
|---|---|---|---|---|
| Chrome 131 | 12 Nov 2024 | Chromium Dash release schedule | 39.8% | Mostly the 2026-08-22 and 2026-08-23 bursts |
| Chrome 78 | 22 Oct 2019 | Chromium Dash release schedule | 28.0% | Present on 11 of 11 observed days |
| Chrome 89 | 2 Mar 2021 | Chromium Dash release schedule | 8.9% | |
| Chrome 151 | 28 Jul 2026 | Chromium Dash release schedule | 4.9% |
Most page requests ask for a page this site never served
11 days of detailed data · to 23 Aug 2026 · sampled, estimates not talliesShare of page requests answered 404, three windows, each named beside its figure.
Each bar is the share of page requests answered 404 in its window, and a 404 counts because this site answers it with a real HTML page.
Without 22 August 2026, half of the requests for a page here ask for one that does not exist. 23 August 2026 carried a second, smaller burst of such requests, so the table also gives the share without both days. Half surprised me. I expected requests for pages that do not exist to be a background hum, not the loudest thing in the log. If your site is small, I would expect yours to look similar. A 404 share this high says more about the requests arriving than about your own broken links.
| Window | Share answered 404 |
|---|---|
| without 2026-08-22 | 50.5% |
| all eleven days | 62.5% |
| without 2026-08-22 and 2026-08-23 | 41.0% |
Most of the 404 numerator comes from unrecognised clients, so this is a claim about request behaviour, never about who is behind it.
Who actually reads the markdown mirror
10 days, without 2026-08-22 · to 23 Aug 2026 · sampled, estimates not talliesShare of markdown mirror fetches by classified client bucket. My own fetches are excluded.
Every page here has a markdown twin at the same address ending in .md, and each bar is one bucket's share of the fetches of those twins.
ClaudeBot is the wrinkle worth naming. In this window the markdown mirror made up 45.7% of its own content fetches, so nearly half of what it took from this site was the mirror, not the HTML. That one I did not see coming. I built the mirror for readers like it and still expected the HTML to win by far more than it did. If you publish a plain-text twin of your pages, the same split in your own logs is what would tell you whether it earned its keep.
Shares describe observed readers, not appetite, because a client that never found the mirror could not have read it, and the AI slice rests on few enough fetches that I hold it at first-signals strength.
| Bucket | Share of markdown fetches |
|---|---|
| SEO tool crawlers | 50.0% |
| Unrecognised clients, not verifiable as people why this is not called users | 18.1% |
| Search engine crawlers | 16.2% |
| AI crawlers | 15.3% |
| Other recognised bots | 0.5% |
I publish a markdown mirror of this page too, at /visibility.md, so a later window can include the readers of the page you are reading now.
Who checks robots.txt before taking content?
10 days, without 2026-08-22 · before the 2026-08-24 policy changeRobots.txt fetches per unit of content taken, both denominators shown for every crawler.
Both columns divide a crawler's robots.txt fetches in this window, first by the HTML pages it took, then by its content fetches with the markdown mirror counted as content.
| Crawler | Per HTML page takenrobots.txt fetches ÷ pages | Per content fetchmarkdown counted as content |
|---|---|---|
| Applebot | 4.0 | 0.8 |
| ClaudeBot | 3.7 | 2.0 |
| OAI-SearchBot | 1.5 | 1.4 |
| Googlebot | 0.7 | 0.6 |
| bingbot | 0.1 | 0.1 |
| PerplexityBot | 0.1 | 0.1 |
| GoogleOther | No robots.txt fetch observed in 10 days. | |
| GPTBot | No robots.txt fetch observed in 10 days. | |
| ChatGPT-User | No robots.txt fetch observed in 10 days. | |
| Amazonbot | No robots.txt fetch observed in 10 days. | |
| YandexBot | No robots.txt fetch observed in 10 days. | |
| CCBot | No robots.txt fetch observed in 10 days. | |
| Baiduspider | No robots.txt fetch observed in 10 days. | |
Applebot, ClaudeBot and OAI-SearchBot fetched robots.txt more often than they took a page, and for 7 of the 13 names no robots.txt fetch was observed in 10 days.
The zeros are the rows I read most carefully. No robots.txt fetch observed in 10 days is not proof of no fetch, because a crawler that read the file once, before the window opened, would look the same here. The ratios are the part you can use, since they say which names check often before taking content from a site like yours.
Every name is self-identified, and etiquette observed is not obedience proven.
Google named the search query behind 2 of 245 impressions
21 days · to 18 Aug 2026 · Google Search ConsoleImpressions whose query Google named, over all impressions, counting only days Google has stopped revising.
Across 21 settled days, Google named the search query behind 2 of 245 impressions, which is 0.8%. The rest are queries Google declined to name. An impression is one appearance of this site in someone's results. A day only counts here once Search Console has stopped revising its figures. That is what I call settled.
The withholding is a known mechanism, not something odd about this site. Ahrefs measured it at 46.77% across the sites it studied, so a site this small sits at the far end of the same mechanism, where almost all of the queries disappear. What I watch is how the share moves, so the table gains a row at each refresh. If your site is small too, the movement in your named share is the part worth keeping, not the level.
| Window end | Settled days | Named-query impressions | All impressions | Share |
|---|---|---|---|---|
| 18 Aug 2026 | 21 | 2 | 245 | 0.8% |
First signals, and what I am holding back
First signals are observations too few to claim a pattern, published with their size stated so nobody mistakes them for one. At launch, exactly one figure on this page is held at that strength, the AI slice of the markdown mirror's readers above. Nothing else here is built on it.
Each finding below waits on the condition beside it, and leaves this list only through a review round.
- llms.txt before and afterThree to four weeks of after-data; the before-baseline and go-live date are already recorded.
- Crawl versus credit, Google editionAll joint Search Console days settled.
- AI training / search / user-fetch splitAI page views in the hundreds; one crawler session currently moves the split by 24 points.
How this is measured
Two archives sit behind this page, and I ask them different questions. The traffic archive holds sampled Cloudflare analytics for this site and tells me what asked for what. The verification archive holds daily counts from checking crawler-name claims against published address ranges and tells me whether a name's addresses matched what its operator publishes. Google Search Console is the third source. It tells me what Google's index did with the site afterwards.
Four rules shape the traffic figures, and you need them to check my work. I count eyeball traffic only, which is Cloudflare's own term for requests arriving from outside its network. A page view is a request for a page-shaped path, meaning any address that is not an asset such as a stylesheet, font or image, answered 2xx or 404, because a 404 on this site returns a real HTML page, and counting successes only would miss most of what actually happens here. I set that rule by checking it against Cloudflare's own page-view total for the same days. My own fetches and this project's tooling come out before anything is counted. Finally, Cloudflare samples this data, so the figures derived from it are estimates, with sampling intervals up to 3.3 observed, and none of them is an exact tally. If you recompute a figure and land a little off mine, sampling is the first place I would look.
Everything my classification list does not match lands in one bucket, labelled "Unrecognised clients, not verifiable as people". I do not present that bucket as human traffic in any chart, table or download. The reason is in the section on what that traffic claims to be.
A client name on this page is the name the client gave itself. A name is not proof of who sent the request, and no operator is accused of anything here.
Since the verification collector first ran on 24 August 2026, I compare any request claiming a known crawler's name, as it is collected, against the address ranges that crawler's operator publishes. The comparison happens in memory. Only counts are stored, and the addresses are discarded. I snapshot the ranges the same day, so when you read a verdict it is against exactly what the operator published then, not what it publishes now.
Five verdict words appear on this page, and they are not interchangeable. Verified means the requests came from inside the operator's published ranges. Not verified means the operator publishes ranges and the requests came from outside them. Unverifiable means the operator publishes no ranges at all, so no check is possible. That is a fact about the operator's documentation, never a verdict on the crawler, and it is why Amazon's Amazonbot and Anthropic's ClaudeBot can show no number. Not covered means the day sits outside verification's reach. Not a user agent means the name is documented as never being an HTTP user agent at all, so no address range could redeem it.
Verification cannot prove intent. It shows that an address sat inside or outside a published range, and that part is measured. Calling the rest a costume carries the same limit as the verdict panel above, and I keep the two apart everywhere on this page. The check also reaches back only seven days from its first run, so days before that reach are unverified permanently. No refresh may add a verdict to an earlier day, and no refresh may turn the reading into a proven claim, whatever later days show.
Some things I deliberately do not collect or publish anywhere on this site. IP addresses. Query strings. Referrers. Bot scores. Raw request paths. Raw user-agent strings. A count of distinct addresses is a count, and it is the only address-shaped thing I will ever publish here.
You will not find visitor counts here. My site is small, and a raw count would make you judge the size instead of the pattern, which is the part you can use. An event count appears only as method context inside one named episode, with its denominator and its window in the same block.
Search Console figures count settled days only. Google keeps revising its most recent days, so I leave an unsettled day out entirely instead of showing it provisionally.
Days that look odd have public explanations. 22 August 2026 carries the wave hour. 23 August 2026 carried a second burst of requests for pages that do not exist, which is why the 404 table names a window without both days. And on 24 August 2026 I changed the site's robots policy, added a tripwire behind it with a block for Bytespider, and took /llms.txt live, so the launch window ends the day before and the etiquette table is a before-baseline. I refresh these figures through the same review round that produced them, and each refresh adds a row to the search series and the after-policy columns to the robots table.
You can download the figures on this page as JSON or CSV and check any number against them. The figures and both files are published under CC BY 4.0. You can reuse them for anything you like, as long as you credit this website. The build cost of this site is measured the same way on the build cost page, and the structured data of the wider web on the structured data reports.
11 days of detailed data, 13 August 2026 to 23 August 2026. 21 settled Search Console days. Built 1 September 2026.
The figures come from this site's own Cloudflare zone analytics and from Google Search Console, read through collectors I run myself. Chrome release dates come from the Chromium Dash schedule service.