web-archive
Archive-It (Internet Archive)
The Internet Archive's subscription archiving service, used by more than a thousand libraries, governments, and universities to build curated web collections that together hold tens of billions of documents. Captures are viewed through replay links on wayback.archive-it.org, organized by collection. It pays off when a specific institution's collection covers your subject; coverage is collection-by-collection, so it is not the place to look up arbitrary URLs.
No programmatic check, opens the archive’s own search.
Why it’s useful & how it works
Archive-It is where institutions such as universities, government bodies, and libraries build and maintain their own focused web collections, so it excels when your subject falls within a specific organization's collecting area. Over 1,000 partner institutions worldwide have created more than 10,000 public collections, all of which are full-text searchable on archive-it.org. You can browse by organization or topic, then view captures through a Wayback-style replay interface. Because coverage is collection-by-collection, it works best when you have identified a relevant collection rather than when searching for an arbitrary URL.
What’s inside
Partners have archived tens of billions of documents collectively. The service has been growing since 2006 and the combined holdings span petabyte scale, though a single aggregate figure is not published separately from the broader Internet Archive.
API access
Per-collection CDX https://wayback.archive-it.org/ <collId>/timemap/cdx?url= (works) ; replay https://wayback.archive-it.org/ <collId>/<ts>/<url> . Aggregate /all/ CDX is blocked.
What we measured
Our own probes, not the archive’s own claims. Re-run periodically; every reading below is dated.
Reachability
- Direct request
- Responded HTTP 200 574 ms
- Through a datacenter proxy
- Responded HTTP 200 1.1 s
- All CDX, direct
- Blocked HTTP 403 325 ms
- All CDX, through a proxy
- Blocked HTTP 403 927 ms
- Collection 2950 CDX, direct
- Responded HTTP 200 323 ms
- Collection 2950 CDX, through a proxy
- Responded HTTP 200 1.1 s
- Collection 2950 replay, direct
- Responded HTTP 200 828 ms
- Collection 2950 replay, through a proxy
- Responded HTTP 200 1.6 s
Did a real search return anything?
Inconclusive · median 840 ms · stable across runs
The per-collection CDX endpoint now answers with a "Session Verification" interstitial rather than JSON, so nothing can be counted. Coverage was always collection-specific and empty for arbitrary URLs: treat as link-only unless a collection is known.
Reachability measured 2026-08-22. Search checked 2026-08-22.
Access
Programmatic API access (a key may be required, see the API tag).