web-archive
Wayback Machine (Internet Archive)
The Internet Archive's flagship web-page archive and the largest of its kind, holding more than a trillion captures of URLs as they appeared over time. Enter any web address to browse its snapshot history in the web interface; free APIs also let you check whether a page is archived and list every capture. It is the default first stop when you need an old or deleted version of almost any web page. The service has had intermittent outages in 2026 and throttles heavy use, though existing captures remain readable.
Why it’s useful & how it works
Reach for the Wayback Machine first whenever you need an old or deleted version of almost any public web page. Enter a URL on web.archive.org and a calendar shows every date a snapshot was taken; pick a date to browse the page as it looked that day, with working internal links. It also offers a free keyless API for checking whether a URL is archived and for listing all captures of an address, making it straightforward to integrate into research workflows. Note that some news publishers have restricted new crawling, though this does not affect existing captures.
What’s inside
Over one trillion web pages captured since 1996, a milestone reached in October 2025. The collection spans roughly 99 petabytes of unique data, expanding to over 212 petabytes with backups, and grows by hundreds of millions of pages per day.
API access
Availability https://archive.org/wayback/available?url= ; CDX http://web.archive.org/cdx/search/cdx?url=&output=json ; TimeMap http://web.archive.org/web/timemap/link/ ; replay https://web.archive.org/web/ <ts>/<url>
What we measured
Our own probes, not the archive’s own claims. Re-run periodically; every reading below is dated.
Reachability
- Direct request
- Responded HTTP 200 1.3 s
- Through a datacenter proxy
- Responded HTTP 200 4.9 s
- API availability, direct
- Responded HTTP 200 787 ms
- API availability, through a proxy
- Responded HTTP 200 702 ms
- API CDX, direct
- Timed out 15 s (This operation was aborted)
- API CDX, through a proxy
- Responded HTTP 200 9.1 s
- API timemap, direct
- Responded HTTP 200 1.3 s
- API timemap, through a proxy
- Responded HTTP 200 3.2 s
Did a real search return anything?
Returned results · 21 results for the sample query · median 25.8 s · flaky across runs
Slow: budget for long waits, or call it asynchronously.
CDX is slow (17-25s); use the Availability API for existence checks, CDX only for full capture lists.
Sample query: http://example.com:80/
Benchmark coverage
Found 9 of 10 test items (90%) in the Web pages / sites cluster. Median response 9.9 s.
By how well known the item is: 3/3 well-known, 3/4 mid, 3/3 obscure .
Reachability measured 2026-08-22. Search checked 2026-08-22. Benchmark run 2026-08-22.
Access
Freely reachable, no key, login, or captcha.