# How these figures are made The five sources behind this site's Schema.org reports, what each one licenses, the rules the pipeline enforces before anything is published, and three mistakes I found in my own published figures and fixed. > Source: https://eduarddziak.com/structured-data/method/ > Snapshot: July 2026 > Updated: 2026-08-19 > Data: /data/schema-stats.json, /data/schema-stats.csv and /data/explorer.json > Licence: Apache 2.0, credit requested. Definitions CC BY-SA 3.0 ## Summary | Measure | Value | | --- | --- | | Real terms measured | 2,487 | | Types where two sources meet | 473 | | Crawl slices compared | 4 | | Snapshot | July 2026 | ## Where two exact counts meet Google and the Web Almanac are the only two sources here that both publish an exact-sounding number for the same term, so this is the one place I can check one measurement against another. Google sorts the terms into six bands and publishes no count inside a band. The Web Almanac publishes an exact count of the sites it found carrying a term, but only within one of four crawl slices, a mobile home page, a desktop home page, a mobile secondary page, or a desktop secondary page. When a type appears in both, I can ask a narrow question. Does the Web Almanac's exact count actually land inside the band Google says it should? Rarely. I found 473 real types that appear in both sources for the newest archived month. I compare them on one crawl slice at a time. Summing the Web Almanac's four slices together would count the same domain more than once, if it carried the type on more than one kind of page. None of the four slices agrees with Google's band more than 15.05% of the time, and the worst of the four manages barely 9.66%. I expected disagreement. I did not expect this much. | Crawl slice | Types compared | Fell inside Google's band | Share inside | | --- | --- | --- | --- | | Mobile home page | 432 | 65 | 15.05% | | Desktop home page | 417 | 48 | 11.51% | | Mobile secondary page | 439 | 50 | 11.39% | | Desktop secondary page | 435 | 42 | 9.66% | When the two disagree, it is almost always the Web Almanac sitting below Google's band, not above it. Google counts across its whole index. The Web Almanac counts a sample of about 15.4 million mobile origins and 12.2 million desktop ones, one home page and one secondary page each, so most of a site's other pages, and most sites outside the sample entirely, are simply not there to be counted. That gap is sample size, not an error in either source, and no correction factor turns one number into the other. I compare the two here. I do not combine them. A blended figure would look more confident than either source really is, and it would be a number you could not check against a download, because no download would agree with it either. So if you ever quote a Web Almanac count for a type, quote it as a count from that crawl and not as the web's total, because the total sits higher and nobody publishes it. ## The five sources, and what I may publish These figures come from five sources, and they measure or document different pieces of the web this project reports on. They license their data on their own terms. The table below names them, what they license, and whether I may redistribute their figures at all. I read the licence fields straight from the same dataset the reports use, so it cannot say something different here than it says on the page where the figures actually appear. | Source | Population | Last measured | Licence | May I redistribute it | | --- | --- | --- | --- | --- | | Google and the Schema.org community, Schema.org usage statistics dataset | Websites in Google's index. Each site is counted once however many pages carry the markup, so this measures adoption by site, not volume of markup. | July 2026 | No licence statement of its own. Both defensible readings — Apache 2.0 as repository content, or CC BY-SA 3.0 as site data — permit archiving, republishing and commercial use with credit. | Yes, with credit | | HTTP Archive Web Almanac 2025, SEO chapter results | About 15.4 million mobile and 12.2 million desktop origins from the Chrome UX Report list, home page and one secondary page each. | 1 July 2025 | Apache License 2.0. | Yes, with credit | | Google Search Central rich result requirements | Not a measurement of the web. What Google requires and recommends per rich result feature. | 11 August 2026 | CC BY 4.0. | Yes, with credit | | Schema.org dated release snapshots | Not a measurement of the web. These are the vocabulary itself, used only to date when a term was invented. | 11 August 2026 | CC BY-SA 3.0 for the vocabulary and documentation. | Yes, with credit | | Web Data Commons structured data class statistics | Pay-level domains in a Common Crawl corpus. Ten yearly releases, November 2015 to December 2024; the series has stalled. | December 2024 | None. Only the extraction software is licensed. The data and workbooks carry no licence at all. | No, cite and link only | The same licence named above for Schema.org's dated releases also covers something more concrete, the description column you read in the explorer table. I reproduce Schema.org's own definitions there. A claim that a term is barely used is worth much less than Schema.org's own words for what the term means, with a link so you can check it yourself. Reusing that one column carries the same share-alike condition Schema.org attaches to it. Reusing anything else on this site does not, because everything else here is my own analysis, not a copy of theirs. ## The rules the pipeline enforces Before any figure reaches a page, a pipeline runs nine numbered checks and refuses to publish if one of them fails. Eight of them write their own sentence into the dataset each time they run, quoted below exactly as generated, ending with the ninth, in the last row. The missing number is the second. It said Web Data Commons' counts and Google's bands must not be combined into one figure, and once the Web Almanac arrived as a third source of exact counts, that rule was folded into the eighth, which now says the same thing for any two of the three. | Rule | What it enforces | | --- | --- | | Rule 1 | Annotation rows removed before any figure was calculated. | | Rule 3 | Terms that cannot be dated are excluded from the adoption-lag figures; 1610 of 2487 were excluded. | | Rule 4 | A term that changes direction between months is reported as unstable, not as movement. | | Rule 5 | Every figure carries the snapshot it came from. | | Rule 6 | Vocabulary counts use Schema.org's own terms only: 937 classes and 1529 properties. | | Rule 7 | Web Almanac type strings were filtered against the vocabulary before joining. | | Rule 8 | The measurements sit in separate blocks and no figure draws on more than one. | | Rule 9 | Every rich result requirement carries its page address and that page's last-updated date. | ## Traps worth knowing about A few numbers in this dataset look like they answer the same question and do not. The table below lists the ones most likely to catch you out, with what is actually true beside it. I have fallen into two of them myself. | Trap | What is actually true | | --- | --- | | Using 5,545 as the denominator for any share of terms. | 5,545 is rows in Google's raw file. 3,058 of them are annotation markers, not terms. The real base is 2,487. | | Reading a Web Almanac count as if it confirms or corrects a Google band. | They are two different measurements of two different populations. Compare them, as the table above does, and never blend them into one figure. | | Treating a missing Web Almanac slice as a count of zero. | 41 of the 473 real types carry no mobile home page figure at all. A missing measurement is not a measured zero. | | Adding the Web Almanac's four crawl slices together for one total. | A domain that carries a type on both its home page and a secondary page would be counted twice. Every figure here stays inside the one slice it was measured on. | | Mixing up the 1,589 undatable terms with the 1,610 excluded from the adoption chart. | 1,589 predates the archive itself, classes and properties only. 1,610 also excludes enumeration values, so it is the one that adds back to 2,487 with the 877 datable terms. | | Treating the 282 rows of other vocabularies as part of the 682 distinct type strings. | 682 counts distinct type strings. 282 counts spreadsheet rows across four crawl slices, a different question with a different denominator. | | Writing a Google band such as 10M+ as if it were a number. | It is a range. Google never publishes the count inside it, here or anywhere else. | ## What the archive keeps I keep the monthly files exactly as Google published them, before I touch anything, so a figure on this site can always be checked against what Google actually said that month and not against a page that may since have changed. Three months are archived so far, and the count grows by one each time I run the collector. Web Data Commons stopped publishing after its December 2024 release, so its own dateline simply ends there. I do not extend it or guess at it. ## Where I have already been wrong Three mistakes reached a live page during this project. All three are worth naming, not folding quietly into a changelog nobody reads. My first release of the flagship report mislabelled a property count. Google asks for a property once per feature, so a property that nine features all ask for was counted nine separate times. I added those counts together, called the total 314, and published 109 of them as thinly used. Counting each property once instead gives a different picture, 206 properties present in the usage file and 103 of them thinly used, a share of 50.00%. The corrected figure is larger, not smaller. The same report also carried a caveat saying some rows were enumeration values rather than properties, and had not been stripped out. I went back to check, expecting to find some. There was nothing to strip. The caveat was true about the risk and wrong about the fact, so I replaced it with the honest one, and I checked thousands of the report's own numbers before and after to make sure nothing else had shifted while I was in there. A field counting live features included one Google is already phasing out, so it read 33 when 32 features are actually live. That single number fed a stat row on two separate report pages, so both were wrong together until I split it into two fields, one for all the features captured and one for only the live ones, both named for what they actually count. All three are now guarded by checks that stop the build if they happen again, and I proved the checks by deliberately breaking them and watching the build fail. That is not the same claim as a clean record. It is the reason to trust the figures here at all, the checks and the corrections, not an absence of mistakes I have not found yet. If you find a fourth, it gets the same treatment. ## The downloads The figures on this site are also files you can open yourself. The dataset behind all the reports, the term-by-term table behind the explorer, and a spreadsheet version of the whole dataset are all linked below, built from the same pipeline run as the pages themselves. | File | What it holds | | --- | --- | | /data/schema-stats.json | The full dataset behind all the reports, Apache 2.0, credit requested. | | /data/schema-stats.csv | The same dataset as a spreadsheet, Apache 2.0, credit requested. | | /data/explorer.json | One row per term, the table behind the explorer. Apache 2.0, except the description column, which is Schema.org's own text under CC BY-SA 3.0. | ## How this page itself is built I built this page from the same two files the other reports read, so the sources table, the rule list, and the comparison above update on their own the next time I run the pipeline. I retype nothing here by hand. That is also the whole reason the rest of this site is worth reading. Analysis © Eduard Dziak, licensed under the Apache License 2.0. Please credit eduarddziak.com with a link. Source data: Google and the Schema.org community, Schema.org usage statistics dataset. Rich result requirements: Google Search Central, licensed CC BY 4.0. Term descriptions: Schema.org, licensed CC BY-SA 3.0, each term linking to its own page. Crawl counts: HTTP Archive Web Almanac, licensed Apache License 2.0.