Where this data comes from

A public monthly crawl of the real web, running since 2011.

You do not have to take our word for any of this. The crawl underneath is open infrastructure — you can query it yourself.

HTTP Archive

An Internet Archive project, supported by Google, Mozilla, Fastly and Akamai. Every month it loads millions of real websites in a real browser and records what it finds, then publishes the whole dataset openly. It has been doing this since 2011.

How a switch is found

01
A real browser loads the page

Not a script scraping HTML — a full browser render, so technologies that only appear after JavaScript runs are still seen.

02
Detected technologies are recorded

Roughly 17.3 per site on average, with version strings where the page exposes them. Across 4,015 technologies.

03
This month is compared against last month

A technology present in one crawl and absent in the next is a candidate change. This is the part nobody else sells well, because it only works if you hold the history.

04
Crawl artefacts are filtered out

If a site times out or blocks the crawler, every technology on it looks dropped at once. We require the domain to be successfully crawled in both months, suppress domains losing an implausible number of technologies together, and require a change to persist into the following crawl before we call it confirmed.

05
Companies are attached

Domains are resolved to registrable form and joined to firmographics — name, size, industry, location — so a change becomes an account rather than a URL.

What that gives you

716,977
domains profiled in depth
4,015
technologies observed
2016
our history begins
monthly
crawl cadence

How far back our observations actually go

HTTP Archive has crawled since 2011, but we read its mobile crawl, and that was a few thousand pages a month until 2018. Our adoption dates begin in 2016, thin until 2018, and are dense from 2021. So a first-seen date is a bound, not a birthday: a company running a technology since 2009 reads 2016 at the earliest, and one that entered the crawl late looks like a recent adopter regardless. Every tenure figure ships with the number of crawls behind it for that reason.

What this method cannot see

Anything invisible to a rendered page: databases, warehouses, server-side infrastructure. Sites with too little traffic to be included in the crawl. Technology behind a login, or on subdomains that are not crawled — which is why coverage thins at the very largest companies, who run their stacks on subdomains. The crawl is monthly and publishes a few weeks after the month it covers, so this is not a same-week trigger. Full limits →

Build a listCoverage & limits