Where the technology data comes from, and how a change is confirmed.
Detection comes from HTTP Archive, an Internet Archive project supported by Google, Mozilla, Fastly and Akamai. Every month it loads millions of real websites in a real browser and publishes what it finds, openly. It has done so since 2011.
A technology present in one crawl and absent in the next is a candidate change. Crawl artefacts look identical to real departures — if a site times out or blocks the crawler, everything on it appears to vanish at once — so candidates pass four filters before we publish them:
1. the domain was successfully crawled in BOTH months 2. it did not lose an implausible number of technologies at once 3. it did not simultaneously deploy new bot protection 4. the change persisted into the following crawl -> status: confirmed
Anything failing 1–3 is suppressed. Anything that has not yet cleared 4 is returned as provisional and is excluded by default; single-month observations reverse about 9% of the time.
confirmed — observed absent across two consecutive crawls. provisional — one crawl only, opt in deliberately. switched — the technology went and something sharing its category arrived in the same month, so both sides were observed.