Просмотр каталога архивов
Все известные нам источники — фильтруйте по категории, кластеру, API или доступу, затем откройте любой архив для подробностей.
- Wayback Machine (Internet Archive) The Internet Archive's flagship web-page archive and the largest of its kind, holding more than a trillion captures of URLs as they appeared over time. Enter any web address to browse its snapshot history in the web interface; free APIs also let you check whether a page is archived and list every capture. It is the default first stop when you need an old or deleted version of almost any web page. The service has had intermittent outages in 2026 and throttles heavy use, though existing captures remain readable. Web pages / sites API
- archive.today (.ph / .is / .md / .li / .vn / .fo) An on-demand archiver where anyone can submit a URL and get a permanent snapshot; it flattens dynamic pages into static copies and often gets past paywalls, with tens of millions of snapshots accumulated. Existing captures are found via its Memento timemap and opened as direct snapshot links on archive.ph and its mirror domains; saving a new page is done on the website and may require solving a CAPTCHA. Reach for it when a page changed or vanished too recently for other archives, or when a paywalled article needs a readable copy. Web pages / sites API captcha
- Common Crawl (CDX index) A nonprofit project that crawls the open web at petabyte scale and publishes everything it collects, with roughly 2.2 billion pages per crawl and a fresh crawl released regularly. Each crawl has a free, keyless index API that reports whether and when a URL was captured, and the stored page content can then be pulled from the public dataset. It is a series of crawl snapshots rather than an on-demand archive, so it suits checking how a site looked in recent crawl windows or doing bulk web research. Web pages / sites API
- Arquivo.pt (Portuguese Web Archive) Portugal's national web archive, preserving billions of files since 1996, including many non-Portuguese sites. Unusually among web archives it offers full-text search over the preserved pages as well as URL lookup, through both its website and free APIs. Turn to it when you need to search inside archived page text rather than only retrieve a known address, or when researching Portuguese and European web history. Web pages / sites API
- Library of Congress Web Archives The Library of Congress's curated web-archiving program: more than a hundred thematic collections of websites selected by its librarians, at petabyte scale. You browse the collections on loc.gov, and collection and item pages can also be fetched as JSON. It shines when your topic matches one of the curated themes, elections, events, organizations, rather than for looking up arbitrary URLs. Web pages / sites API
- Perma.cc A permanent-archiving service from the Harvard Library Innovation Lab, designed so links cited in legal and academic writing never rot; it holds millions of permalinks. Each archived page gets a short perma.cc code that replays the capture, and a free public API lets you look existing records up. Most useful when following a Perma link from a court document or paper, or when checking whether a page was preserved as a citation. Web pages / sites API
- Archive-It (Internet Archive) The Internet Archive's subscription archiving service, used by more than a thousand libraries, governments, and universities to build curated web collections that together hold tens of billions of documents. Captures are viewed through replay links on wayback.archive-it.org, organized by collection. It pays off when a specific institution's collection covers your subject; coverage is collection-by-collection, so it is not the place to look up arbitrary URLs. Web pages / sites
- UK Government Web Archive (The National Archives) The National Archives' public archive of UK central-government websites, with billions of records reaching back to about 2003. Archived pages are opened through replay links in its Wayback-style viewer on webarchive.nationalarchives.gov.uk. It is the place to find old versions of gov.uk and departmental sites, especially while the main UK Web Archive is offline. The server rejected our automated checks, so behaviour may be inconsistent; use the website directly. Web pages / sites
- Vefsafn.is (Icelandic Web Archive) Iceland's national web archive, run by the National and University Library, which has harvested the entire .is domain about three times a year since 2004 and holds Iceland-related sites back to 1996. Snapshot lists come from its Memento timemap, and individual captures open as replay links on vefsafn.is, where the viewer may ask you to pass a CAPTCHA. The natural source for the history of Icelandic websites. Web pages / sites API captcha
- Trove / Australian Web Archive (NLA) The National Library of Australia's web archive. PANDORA, the Australian Government Web Archive, and .au domain crawls combined, holding billions of files and searched through the Trove portal. Captures open as replay links on webarchive.nla.gov.au; the Trove API does not cover archived websites, so searching happens on the site itself. The primary source for the history of Australian websites. Web pages / sites
- DNB Webarchiv (German National Library) The German National Library's selective web archive, collecting thematic snapshots of German websites since 2012. There is no search API: you browse the thematic collections through the portal, which are free to view worldwide, while deeper content is restricted to the library's reading rooms. Worth checking when tracing German sites or subjects its collections cover.
- Ghostarchive A small independent on-demand archiver, notable for handling YouTube videos and Twitter posts as well as ordinary web pages. Snapshots are reached by direct links to individual archive pages; there is no search API. Check it when hunting for preserved copies of social-media content that mainstream web archives tend to capture poorly. Web pages / sites
- Conifer (Rhizome, ex-Webrecorder.io) A hosted capture service from Rhizome (formerly Webrecorder.io) for high-fidelity, interactive recordings of web pages, gathered into thousands of user-created collections. You browse collections and open individual recordings on the website; there is no public search API. Be aware the service is winding down: capture and editing stop in May 2026 and from June 2026 it becomes read-only, though existing collections stay viewable.
- Webrecorder / Browsertrix The maker of Browsertrix and the successor toolchain to Conifer: self-hostable or hosted crawling software that produces WACZ/WARC web-archive files. There is no central public corpus to search, every deployment holds its own captures. Relevant when you want to create high-fidelity archives of sites yourself rather than search existing ones.
- Stanford Web Archive Portal (SWAP) Stanford Libraries' access portal for its curated web-archive collections, which are stored through the Archive-It service. There is no search API on the portal; you browse and open captures on the website. Mainly useful when Stanford's curated collections happen to cover your subject, the same material is also reachable via Archive-It.
- EU Web Archive (Publications Office) The EU Publications Office's archive of the websites of EU institutions, agencies, and bodies, collected since 2018. It is browsed through the portal, with no documented search API. The place to look for earlier versions of europa.eu and other institutional sites; note that access is blocked from some networks, so it may not load on every connection. IP-blocked
- KB Netherlands Web Archive The Royal Library of the Netherlands' web archive ('Websites van Nederland'), preserving about 25,000 curated Dutch sites, several thousand of which have already disappeared from the live web. There is no public search API, and deeper access to archived content is largely limited to the library's reading rooms. Best used as a pointer that a vanished Dutch site may be preserved, with full consultation done on site. IP-blocked
- US NARA Web Harvests (webharvest.gov) The US National Archives' collection of periodic web harvests of congressional and presidential websites. It is browse-only on webharvest.gov, with no search API. Reach for it when you need to see how official US government sites looked at the close of a Congress or presidential administration.
- Kiwix / Zimit An offline-archive ecosystem rather than a lookup service: Kiwix packages whole websites. Wikipedia, StackExchange, and thousands more, into downloadable ZIM files, and Zimit can convert other sites into the same format. You browse the package library and download complete site copies for offline reading; there is no per-URL snapshot search. Useful when you want a full working copy of a site rather than a dated capture of a single page.
- PADICAT (Catalonia) Catalonia's national web archive, combining curated selections with crawls of the .cat domain. Everything happens through its own web portal, no confirmed API, so you search and view captures on padicat.cat. The go-to source for the history of Catalan websites, though the site can be slow to respond.
- Croatian Web Archive (HAW) The Croatian national web archive, run by the National and University Library in Zagreb, preserving .hr sites through curated selection and annual domain crawls. Access is by browsing the HAW website; no API is documented. It covers Croatian web history, but the site was unreachable in our recent checks, so its availability is currently uncertain.
- Bibliotheca Alexandrina Web Archive The web-archive arm of Egypt's Library of Alexandria, which historically hosted a mirror of the Internet Archive's pre-2007 collection alongside regional material. Access is by browsing the site; there is no API. Its present holdings are unclear, the page loads, but there is no recent evidence the mirror still functions, so verify content before relying on it.
- Yandex (cache / SERP) The Russian search engine, listed here because it is the last major engine still surfacing cached copies of web pages. There is no stable API: you run an ordinary search on yandex.com and look for the 'saved copy' link beside a result. Worth a try when a page disappeared too recently for the web archives to have caught it. Expect anti-bot checks, the site frequently challenges visitors before showing results. Web full-text captcha
- Crossref REST API The authoritative registry of DOI metadata for scholarly publishing, covering roughly 165 million records. Look up any DOI to retrieve its bibliographic record, or search by title and author, on the website or through a free open API. It is the backbone reference for verifying citations and finding out what a DOI actually identifies. Academic papers API
- Unpaywall A database that maps DOIs to legal open-access copies of papers, tracking OA status for around 150 million DOIs. Give it a DOI and it reports whether a free, legitimate full-text copy exists and where, via a simple free API. The standard way to answer 'can I read this paywalled paper for free somewhere legitimate?' API
- OpenAlex An open index of the global research landscape: about 480 million scholarly works linked to authors, institutions, and topics. You can explore it at openalex.org or query its API, which since February 2026 requires a free account key (with a substantial free daily allowance). Strong for broad literature mapping, works, authors, and institutional output in one place. Keyless access is now limited to a small daily demo quota. Academic papers API · key
- Semantic Scholar (Academic Graph) An academic search engine and citation graph from the Allen Institute for AI, covering about 220 million papers with citation context. Search on the website, or use the free Graph API, which needs no key but runs on a shared, heavily throttled anonymous pool. Especially handy for following citation trails and seeing how papers relate. Under load the keyless API is frequently rate-limited, so the website is the more dependable route. Academic papers API
- DataCite REST API The DOI registry for research outputs beyond journal articles, datasets, software, and preprints, covering tens of millions of DOIs. Records can be looked up or searched through a free API with no key required. It is the complement to Crossref when the object you are tracing is a dataset or a piece of research software rather than a paper. Academic papers API
- OpenCitations An open index of scholarly citation links, with more than two billion DOI-to-DOI connections as of its February 2026 refresh. Citation counts, incoming citations, and reference lists for any DOI are available through a free API; the website's search page sits behind a CAPTCHA, but the API answers cleanly. Use it to see who cites a paper or to build citation trails from fully open data. API
- dblp (CS bibliography) The curated computer-science bibliography, indexing around seven million publications from CS journals and conferences with carefully maintained author records. Search it on dblp.org or through its free API. Its scope is strictly computer science, but within that field it is often the cleanest, most reliable record of who published what, where, and when. Academic papers API
- arXiv (export API) The preprint server for physics, mathematics, computer science, and neighboring fields, hosting about 2.5 million papers with free full text. Search and browse on arxiv.org, or query the free export API. It is the first place to look for new research before, or instead of, journal publication. Automated queries have been throttled more tightly since early 2026, so machine access may need to slow down. Academic papers API
- bioRxiv The biology preprint server run by Cold Spring Harbor Laboratory, hosting over 300,000 preprints with free full text. You search and read on biorxiv.org, the site sometimes shows a brief Cloudflare check before loading, and per-paper details are also available through a free API. The place to find biology research months before it appears in journals. Academic papers API captcha
- medRxiv bioRxiv's sister preprint server for medicine and the health sciences, with over 60,000 freely readable preprints. Use the medrxiv.org website to search and read (it may show a brief browser check first); paper details are also exposed through the shared bioRxiv API. Reach for it when tracking clinical and public-health research ahead of peer review. Academic papers API captcha
- PubMed / PMC (NCBI E-utilities) The US National Institutes of Health's index of biomedical literature, covering roughly 37 million citations plus about 10 million full-text articles in PubMed Central. You search the corpus through NCBI's E-utilities, which return results without an API key. It is the standard first stop for medical, clinical, and life-sciences literature. Academic papers API
- Europe PMC A life-sciences literature collection maintained by EMBL-EBI, spanning around 43 million publications along with open full-text content. A single keyless search interface returns both metadata and, where available, the full text of an article. Reach for it when you want biomedical papers and need the actual article text rather than just a citation. Academic papers API
- DOAJ A directory of vetted open-access journals together with article-level metadata, covering roughly 21,000 journals and over 10 million articles. Searches run through a keyless interface over its openly licensed metadata, or you can browse the site directly. Use it to confirm a journal is genuinely open access or to find freely readable articles. Academic papers API
- DOAB (Open Access Books) The book counterpart to DOAJ: a directory of peer-reviewed open-access academic books, numbering around 95,000 titles. Its catalogue is searchable through a keyless interface and is also browsable on the web. Turn to it when you need scholarly monographs and edited volumes that are free to read. Books API
- OpenAIRE Graph A European open-science knowledge graph that links hundreds of millions of research products, publications, datasets, software, and funding records. You can query it through its API or browse the Explore portal on the web. It is a good choice for tracing the connections between a paper and its underlying data, software, or grant, though the service is mid-migration to a new graph API in 2026 and its older search endpoints are being phased out. Academic papers API
- CORE (core.ac.uk) An aggregator of open-access research that harvests full text and metadata from more than 10,000 repositories, totalling some 300 million records with over 40 million full-text items. Access is through a search API that needs a free key, since anonymous use is heavily throttled. Reach for it when you want the widest possible net over open-access full text rather than a single repository. API · key
- BASE (Bielefeld) A Bielefeld University academic search engine indexing over 400 million documents harvested from more than 12,000 content providers. The web portal is openly browsable for searching across this material. Its programmatic interface is limited to pre-registered IP addresses, so in practice you use it through the website. IP-blocked
- The Lens A combined search and analytics platform covering scholarly works alongside patents, with more than 270 million records. The web interface lets you search and cross-reference research and patent literature in one place. Its API is approval-gated and largely paid, so the practical route is the website, useful when a question spans both academic publications and patents.
- Dimensions (Digital Science) A linked research database connecting publications, grants, patents, clinical trials, and policy documents, with over 140 million publications. You search it through its web application. Programmatic access is subscription- and approval-gated, so use the site directly when you need to follow a topic across funding, trials, and policy as well as papers.
- Scite.ai A citation-analysis service that classifies roughly 1.6 billion citations by whether they support, contrast with, or merely mention a claim. You reach a paper's citation report through deep links to its website. The API is enterprise-priced, so access is link-based, helpful when you want to know not just who cited a paper but how. captcha
- Connected Papers A tool that builds a visual graph of papers related to a given work by similarity, drawn from a corpus of tens of millions of papers. You use it by opening the graph view for a paper on its website. There is no open lookup interface, so it serves as a deep link, reach for it to discover adjacent literature around a paper you already know.
- Open Library An open, editable catalogue of books, editions, and authors run by the Internet Archive, holding records for over 40 million editions and works, with links to borrowable copies. You can search by title or look up a specific ISBN through its keyless API, or browse the catalogue on the web. It is a solid general-purpose stop for bibliographic details and for finding where a book can be read or borrowed. Books API
- Internet Archive Texts The text and book collection on archive.org, comprising more than 40 million digitized items, some available only through controlled lending. A keyless search-and-metadata interface lets you find items and open them in the reader. Use it to establish that a book or document exists and to read public-domain material, bearing in mind that in-copyright full text is lending-gated and the service can throttle under heavy load. Books API
- scholar.archive.org / fatcat The Internet Archive's scholarly catalogue (fatcat), offering full-text search across more than 25 million open-access papers, including preservation copies of articles. You search it through the scholar.archive.org website. It is best reached via the web interface for now, since its lookup API was unreliable when last checked. Academic papers
- HathiTrust A large collaborative digital library of over 18 million digitized volumes drawn from research libraries. A keyless bibliographic interface returns catalogue records with deep links into the reader; public-domain volumes are readable while in-copyright full text is restricted to affiliated members. Reach for it to verify holdings and access scanned volumes, though its search interface can present a challenge page from some networks. Books API captcha
- Google Books API Google's index of book volumes, covering tens of millions of titles with metadata and, where permitted, previews. You can search by title or ISBN through its API, and a free key raises the otherwise tight request quota. It offers broad coverage for confirming a book's existence and edition details and for finding preview pages. Books API
- Project Gutenberg (Gutendex) The classic collection of around 75,000 free public-domain ebooks, with the Gutendex service exposing its catalogue as JSON for searching by title or author. You can also browse and download directly from the Gutenberg website. It is the go-to for complete, free ebook texts of out-of-copyright works, though the catalogue query can be slow or occasionally time out. Books API
- Standard Ebooks A small, carefully produced collection of about 900 public-domain ebooks, polished and formatted to a high standard, with an OPDS feed. You browse and download titles from its website. Reach for it when you want a clean, well-typeset edition of a classic rather than a raw scan. Books IP-blocked
- Wikisource Wikimedia's library of free-content source texts, transcribed books, documents, speeches, and historical material, running to millions of pages. You can search it through the standard MediaWiki API or on the website. It is useful for finding transcribed, citable full texts of public-domain and primary-source documents. Books API
- WorldCat / OCLC The world's largest union catalogue of library holdings, aggregating more than 550 million records and showing which libraries hold a given item. You search it through the web interface at search.worldcat.org. There is no longer a free API, so use the website, it remains the reference standard for locating a physical copy in a library.
- Trove books (NLA) The National Library of Australia's discovery service for books, newspapers, images, and Australian material, spanning billions of items. You can search it on the Trove website, or through its API with a free key. It is the principal resource for Australian published and historical material; note this covers books and catalogue records, not web-archive captures. API · key
- Anna's Archive A meta-search index over several shadow libraries. LibGen, Sci-Hub, Z-Library, plus scraped collections, covering roughly 64 million books and 96 million papers. It points to externally hosted copies of works, many of them copyrighted, so it functions as a link directory rather than serving files itself. Researchers use it as a single search across multiple shadow sources, but its domains change frequently and the working mirror must be resolved at the time of the query. Books captcha
- Library Genesis (LibGen) A long-running shadow library of books, textbooks, and articles numbering in the millions. It indexes copyrighted material hosted elsewhere, so it acts as a link source rather than a content host. Its mirror domains change often and must be resolved at query time, which makes access intermittent. Books IP-blocked
- Sci-Hub A service that provides full-text PDFs of paywalled journal articles looked up by DOI, with a corpus of around 88 million papers. It links to copyrighted material hosted elsewhere rather than serving it directly. The collection is largely frozen at around 2021, so it is a best-effort route for older DOIs, and its mirror domains rotate. Academic papers captcha
- Z-Library A large shadow library of ebooks and articles, holding tens of millions of items behind an account login. It indexes copyrighted material hosted elsewhere, so it acts as a link directory, and numerous phishing clones imitate it. Access is fragile because of the login wall and frequently rotating domains. captcha
- Bluesky public AppView (AT Protocol) An open, keyless interface to the whole Bluesky social network, profiles, posts, and feeds across tens of millions of accounts. Public reads of profiles and feeds run through its API without any login or key. It is the cleanest way to pull current social posts and profiles directly into search results. API
- Telegram public preview (t.me/s/) Telegram's own server-rendered web preview of any public channel, showing recent posts without an app or login. You reach it simply by opening the channel's preview page on t.me. It is the most straightforward way to read and pull posts from public Telegram channels. API
- Wayback, twitter/x URLs The Wayback Machine applied to twitter.com and x.com URLs, returning archived snapshots of tweets and profiles where they were crawled. You can check whether a specific tweet was captured and open the saved copy via replay links. It is among the more reliable ways to recover deleted tweets, though coverage is patchy and strongest for high-profile accounts. Social API
- Wayback Tweets (claromes) A tool that queries the Wayback Machine's capture index to list every archived tweet for a given handle. It draws on the same underlying Wayback capture data, scoped to one account at a time. Reach for it to enumerate a user's surviving archived tweets, though the hosted app itself often sleeps or times out. Social API
- archive.today. X / Facebook archive.today used specifically for X and Facebook pages that the Wayback crawler often cannot reach, holding tens of millions of on-demand snapshots. Existence is checked through its Memento timemap, and you open the saved snapshot via a direct link. It is valuable for social pages other archives miss, though the snapshot pages can rate-limit and it has faced regional blocks in 2026. Social API captcha
- Wayback, facebook.com The Wayback Machine applied to facebook.com URLs, surfacing archived snapshots of public Facebook pages and posts where they were captured. You check capture existence and open saved copies through its replay links. Coverage is sparse because Facebook blocks crawlers and many captures are only the login interstitial, so results are best for public Pages. Social API
- Wayback, reddit.com The Wayback Machine applied to reddit.com URLs, holding archived snapshots of threads and profiles with good coverage of popular discussions. You check whether a thread was captured and open the saved copy via replay links. Because it does not depend on Reddit's own API, it is one of the more reliable ways to recover removed or deleted Reddit content. Social API
- IA Twitter Stream Grab A set of Internet Archive bulk dumps of the old Twitter public sample stream, stored as monthly JSON files and covering billions of tweets from roughly 2011 to 2023. You access it by downloading the dataset files, which are listed via the item's metadata API. It is meant for bulk, offline analysis of historical Twitter data rather than looking up an individual tweet; the dataset is frozen, since the source stream no longer exists. Social API
- PullPush.io An open successor to Pushshift offering keyless full-text search across Reddit comments and submissions, including content that was later removed or deleted. You query it through its JSON API by subreddit or keyword. It is the main practical tool for searching Reddit's history and recovering removed posts. Social API
- Arctic Shift A Reddit data service providing a keyless API, downloadable bulk dumps, and a web interface, with more generous limits than PullPush. Its full-text search is scoped to a specific user or subreddit rather than the whole site. Reach for it as a complement to PullPush when you need archived Reddit data within a particular community or account. Social API
- Reddit Data API (official) Reddit's own data API, where appending .json to a page URL returns its content without authentication and OAuth unlocks more. It serves live, current Reddit data, posts, comments, and profiles, but offers no deep historical search. Use it for present-day Reddit content and pair it with PullPush or Arctic Shift when you need the past. API
- Desuarchive (4chan archive) A FoolFuuka-based archive of selected 4chan boards (/a/, /int/, /k/, /tg/, /wsg/ and others) with a large history. It keeps full-text-searchable copies of threads, so you can find old 4chan posts after they vanish from 4chan itself, either on the website or through its keyless JSON API. A core source for forum research on the boards it covers; note that some boards carry NSFW or extreme content. Social API
- Mastodon per-instance API Every Mastodon server exposes a built-in REST API, and public reads such as account lookups and public timelines generally work without a token. There is no global search: each query targets one specific instance, federation gives only partial cross-instance reach, and full-text search requires a token on most servers. Reach for it to check a particular Mastodon account or browse an instance's public posts. API
- Fediverse Observer A directory and uptime monitor covering thousands of Fediverse/Mastodon servers, with a GraphQL endpoint behind it. Use the website to discover which instances exist and which are currently alive. It holds no posts itself, it is the natural first stop before digging into a specific Mastodon server.
- Nitter (xcancel.com et al.) Privacy front-ends for X/Twitter (xcancel.com and similar Nitter instances) that show profiles and timelines without logging in, including per-user RSS feeds. They display live X data only, they are not an archive, and the ecosystem has mostly collapsed, with just a few instances clinging on. Useful for reading an X account without an account of your own, but expect blocks and outages; check the Nitter health tracker for a working instance first. IP-blocked
- Nitter health tracker A live status page tracking which of the remaining Nitter instances are currently up. It contains no content of its own, it is a utility you open first to pick a working Nitter mirror before trying to view X/Twitter profiles through one. Just open it in the browser; there is nothing to search.
- Thread Reader App Unrolls X/Twitter threads into single readable pages, and once a thread has been unrolled the page persists even if the original tweets are deleted. Millions of unrolled threads are readable without any login at predictable /thread/<tweetid>.html addresses. Handy for recovering a deleted thread that someone unrolled while it was live; creating new unrolls requires an X account to mention the bot.
- Politwoops (ProPublica) ProPublica's archive of tweets deleted by politicians, holding around a million deletions collected between 2012 and 2023. You browse it on the web, there is no API, looking up a politician to see what they removed. The collection is frozen: tracking ended in 2023, so it is strictly a historical record of that period.
- Telegago (Google CSE) A Google Custom Search configured to dig through public Telegram (t.me) pages, channels, posts, and profiles that Google has indexed. You use it like an ordinary search box in the browser; coverage is whatever Google happened to crawl of t.me. A quick, free way to discover Telegram channels on a topic when you don't know where to start.
- TGStat The largest catalog and analytics service for Telegram, covering over 2.5 million public channels and groups with post search and audience statistics. Use the website to find channels by topic or keyword and inspect their reach; the API is paid and token-gated. The go-to when mapping the Telegram ecosystem around a subject or researching a specific channel. captcha
- Telemetr.io A large Telegram analytics index, similar in role to TGStat: channel statistics and post data for the public Telegram sphere. Full data sits behind a login and paid subscription, so treat it as a site you open and explore rather than something to query openly. It was unreachable in our last availability check, so expect possible downtime.
- SocialData API A commercial API that returns X/Twitter data, search, user profiles, and tweets, without going through the official X API. It is purely programmatic: requests need a paid API key, billed per tweet retrieved, and there is no free browsing interface. Relevant when you need dependable X search results in volume and have a budget for it. API · key
- X API v2 (official) The official developer interface to X (Twitter), covering the full platform. Since early 2026 it runs on a pay-per-use credit model with the free tier eliminated, and full-archive search is gated to expensive high tiers. For archive research it is now a last resort, practical only if you are prepared to pay per request, this site links out to it rather than querying it.
- Meta Content Library Meta's research tool for studying public Facebook and Instagram content, the successor to CrowdTangle, exposing posts with over 100 data fields. Access is restricted to vetted academic and nonprofit researchers, who work inside Meta's secure environment after ICPSR vetting; there is no public website or open API. Listed so researchers know the official route into Facebook/Instagram data exists; apply through ICPSR if you qualify. IP-blocked
- Reveddit Shows Reddit content removed by moderators or AutoModerator for a given user or thread. You use it on the web by entering a username or thread link; it has no stable API. Since losing its Pushshift data source its depth is much reduced, mainly mod-removed content, and the site sits behind an anti-bot challenge, so treat results as best-effort. captcha
- redditsearch.io An older Reddit search interface that was powered by the Pushshift archive. With Pushshift's public access gone it is effectively non-functional for search, even though the page still loads. Kept only for completeness, for historical Reddit content use PullPush or Arctic Shift instead.
- twstalker / Sotwe (X viewers) Third-party viewers (twstalker, Sotwe and similar) that scrape X/Twitter and re-display a profile's recent tweets without login. Coverage is per-profile and recent posts only, accessed by opening the site in a browser. These scrapers are fragile and frequently blocked, our checks hit errors and anti-bot challenges, so treat them as an occasional fallback, not a dependable source. captcha
- Picuki (IG viewer) An anonymous Instagram viewer for browsing public profiles, posts, and stories without an account. It works purely as a website, with no API, in a constant tug-of-war with Instagram's defenses. Useful for a quick no-login look at a public IG account, but it sits behind anti-bot challenges and breaks intermittently. captcha
- Imginn (IG viewer) A no-login viewer and downloader for public Instagram posts and stories. It is browser-only with no API, and like other IG viewers it is scrape-based and fragile. Our availability checks found it blocked or challenged, so keep alternatives ready if it doesn't load. captcha
- AnonyIG (IG story viewer) One of a class of anonymous Instagram story and highlight viewers that require no login. You use it directly in the browser to look at a public account's stories; there is no API. Instagram actively breaks these services, our checks met legal blocks and errors, so expect to try several similar viewers before one works. IP-blocked
- Bluesky Firehose / Jetstream Bluesky's real-time event stream covering the whole network as it happens; Jetstream is the companion service that delivers the same stream as simpler JSON. It is a keyless developer-facing WebSocket feed, built for live monitoring and data collection rather than looking up old posts. Relevant when you want to capture or watch Bluesky activity in real time. API
- 4plebs (4chan archive) A long-running archive of several 4chan boards (/pol/, /x/, /tv/, /adv/ and others) holding years of full-text-searchable history. Search it through the website; its JSON API exists but anti-scraping defenses block it from many networks, so automated access is unreliable. A key source for tracing old or deleted 4chan threads on its boards, be aware it carries NSFW and extreme content. Social IP-blocked
- archived.moe (4chan) A 4chan archive notable for unusually broad multi-board coverage and a large history. Plan on using the website directly: an API technically exists but anti-bot defenses block automated clients at both the pages and the API. Worth checking when a thread isn't on Desuarchive or 4plebs, though its anti-scraper protections can challenge ordinary visits too; not to be confused with the long-dead archive.moe. captcha
- Software Heritage A universal archive of source code, the 'Library of Alexandria of code', preserving over 20 billion unique source files from more than 350 million projects. You can search for projects, look up files by hash (sha1/sha256/git), and resolve permanent SWHID identifiers, on the website or through its keyless REST API. The place to go when a repository has vanished from its original host, or when you need to identify a file by its hash. Software / packages API
- Wikimedia Commons The largest repository of freely licensed media, over 110 million files, behind Wikipedia and its sister projects. Everything is searchable on the website, and the standard MediaWiki API serves the same search without a key. Reach for it when you need openly licensed images and other media with documented sources and licensing. Images API
- Openverse A search engine from the WordPress project spanning more than 800 million openly licensed and public-domain images and audio files drawn from many source collections. Search on the website or via its keyless API. Best when you want reusable media and one query across many open collections instead of visiting each separately. Images API
- Internet Archive (Video) The Internet Archive's video holdings: millions of items spanning films, television, and ephemera. Search and metadata are available through the website or the keyless archive.org search API. A primary source for old, obscure, or otherwise vanished footage; it shares the Internet Archive's intermittent 2026 outages and rate-limiting, so occasional retries may be needed. Video / audio / music API
- Internet Archive (Audio) / Live Music Millions of open audio items at the Internet Archive, including the Live Music Archive of over 250,000 legally taped concerts. It is searchable on the website, with the same keyless search and metadata API as the rest of archive.org. The first stop for live recordings, old broadcasts, and other audio that is hard to find anywhere else. Video / audio / music API
- YouTube via Wayback Machine A recovery technique rather than a separate site: looking up old captures of YouTube watch pages in the Wayback Machine to retrieve the metadata, titles, and thumbnails of deleted or changed videos. You check whether a capture exists via the Wayback availability API, then open the archived page itself. Coverage is limited to the watch URLs the Wayback Machine happened to save, and you usually recover the page rather than the playable stream, pair it with Filmot or Ghostarchive when hunting a lost video. API
- Dailymotion A major Western video host with hundreds of millions of videos and notably stable search through its keyless public Graph API. In our research on where vanished YouTube videos resurface, it stood out as the best openly accessible reupload source. Search it when chasing a copy of a video that has disappeared from YouTube; here it is currently offered as a link out to the site. Video / audio / music
- RuTube Russia's main video host, with tens of millions of videos; our research found roughly a 4% overlap with YouTube reuploads. Search works on the website, though its informal search endpoint rate-limits and throws captchas at bursts of automated traffic. Worth checking when a Russian-language or Russia-related video has vanished from YouTube. Video / audio / music captcha
- VK Video The video arm of VKontakte, holding billions of clips with high potential for finding Russian-language reuploads. Most of it sits behind a login wall: web search requires an account, and the official API needs a registered app and access token. Check it when chasing RU-sphere video content, but expect to sign in to search. IP-blocked
- Odysee A video platform built on the LBRY protocol, hosting tens of millions of items with notably open access, no captchas or login walls in our research. You can search it directly on the website. A solid reupload candidate to check when a video has been removed from mainstream platforms. Video / audio / music
- Rumble A US video host with hundreds of millions of items. It has no stable public API and blocks datacenter traffic, so search it directly on the website from an ordinary connection. Worth including when sweeping video hosts for reuploads of content that has disappeared elsewhere. Video / audio / music IP-blocked
- BitChute An independent video host holding millions of videos. There is no stable public API and the search page is gated behind an hCaptcha against automated visitors, so search it manually in the browser. A standard stop when sweeping alternative video platforms for copies of removed or hard-to-find footage. Video / audio / music captcha
- OK.ru Video The video arm of Odnoklassniki, a major Russian social network hosting hundreds of millions of clips. It is mainly of interest as a place where Russian-language footage or reuploads of vanished videos may survive. There is no open query route, the official API requires a registered app and key, so you use its on-site search via a direct link. Currently degraded: the site returns empty results to automated connections, so search it manually in a browser. IP-blocked
- Bilibili China's largest video platform, hosting hundreds of millions of videos. For an archives researcher it matters as a possible home for reuploads and Chinese-language material that exists nowhere else. There is no keyless English-facing API, and search from outside China is gated by geo-checks and captchas, so this is a follow-the-link resource: open the site and search there. Video / audio / music captcha
- Torrents-CSV An open torrent search engine built on a community-maintained CSV dataset, indexing millions of torrents; the whole service is also self-hostable. It exposes a simple keyless JSON API, so searches run automatically and return clean results. Useful for checking whether a file or release that disappeared from the web still circulates on BitTorrent. It only hands out links to torrents, so weigh the legality of anything you fetch. Torrents API
- Academic Torrents A platform for distributing research datasets and academic data over BitTorrent, petabytes of legitimately shared material. You can search and browse collections freely, and reads need no key or account. Reach for it when a large public dataset is easier to fetch as a torrent than from a slow or fragile institutional server. Torrents API
- The Pirate Bay (apibay) The longest-running general torrent index, with a very large catalogue across all content types. The practical way in is apibay.org, its keyless JSON search backend, since the main website is widely blocked by ISPs. Researchers use it to check whether vanished software, media or documents still circulate as torrents. Results point largely to copyrighted material, treat them as links only. Torrents API
- Nyaa The leading torrent index for anime and East Asian media. Its RSS feed doubles as a simple keyless query API, and results link out to the torrent pages. The place to look for old fansubs, out-of-print releases and Japanese media that never got a Western edition. As with any torrent index, treat results as links of varying legality. Torrents API
- SauceNAO The de-facto reverse-image search for anime, manga and fan art, indexing billions of images from sources such as Pixiv and Danbooru. Submit an image or image URL and it identifies where the artwork came from; there is a JSON API, but it requires a free registered key. Indispensable when you need to trace an illustration back to its original artist or first posting. Images API · key
- VirusTotal The canonical file-by-hash lookup: it aggregates antivirus verdicts and sandbox analyses for billions of files, URLs and domains. Paste a hash or URL into the website for a consolidated report, or use its JSON API with a free key for automated lookups. Archives researchers reach for it to identify a mystery file or check whether a hash matches known malware before handling it. File by hash API · key
- MalwareBazaar (abuse.ch) A free malware-sample repository from abuse.ch, holding millions of samples searchable by hash, tag or malware signature. Queries go through its API, which now requires a free abuse.ch Auth-Key (one key covers their related services). Use it when you have a suspicious file's hash and want to know whether it is a catalogued sample, or need the sample itself for analysis. API · key
- URLhaus (abuse.ch) abuse.ch's database of malware-distribution URLs, holding millions of malicious links searchable by URL, host or payload hash. Lookups run through a simple API using the same free abuse.ch Auth-Key as their other services. Check it when you want to know whether a URL found in an archived page or old dataset was known to push malware. API · key
- ThreatFox (abuse.ch) An open indicator-of-compromise sharing platform from abuse.ch, cataloguing millions of IPs, domains, hashes and URLs tied to named malware. It is queried via its API with the free abuse.ch Auth-Key. Useful for putting an artifact in context: search a hash or domain and learn which malware family or campaign it has been associated with. API · key
- Triage (tria.ge) An automated malware sandbox run by Recorded Future, with a large public corpus of analyzed samples. Public reports are searchable by file hash, and API access uses a token from a free Researcher account. A good next step after a hash lookup elsewhere, when you want to see what a sample actually does when executed. API · key
- MalShare A community-run public malware repository with millions of samples, offering hash search and sample download. Its straightforward API is keyed but free, self-serve registration grants 2,000 calls a day. A lower-barrier alternative to invite-only repositories when you need to check or retrieve a sample by hash. API · key
- Maltiverse A threat-intelligence aggregator for checking IPs, domains, URLs and file hashes against a large pooled IOC dataset. Lookups go through its JSON API, with keys available on a free tier. Use it to enrich an artifact from your research, a hash or domain pulled from an archived page, say, with what threat feeds collectively know about it. API · key
- Flickr One of the largest photo-sharing archives on the web, holding billions of photos, a deep well of historical and community photography. Search works on the website, and a mature REST API exists for keyed access. Worth searching when you are chasing original photographs of places, events or communities from the earlier web. Degraded for new integrations: fresh API keys now require a paid Flickr Pro subscription (existing keys keep working), so plan on the web interface. API · key
- Imgur A major image host that long served as the default picture backend for Reddit and countless forums. Old uploads can be retrieved by image ID, on the site or via an API that needs a free Client-ID. It mostly matters for resolving imgur links found in archived discussions. Be aware its archival value dropped after the 2023 purge of anonymous content, many old IDs now return nothing, so verify a link still resolves before relying on it. API · key
- DeviantArt A long-standing online art community hosting hundreds of millions of artworks. You can search and browse freely on the site, and an OAuth-based API covers public browsing and search for registered apps. Researchers come here to trace artwork provenance, recover an artist's public gallery, or date a piece of fan art. API · key
- Freesound A collaborative database of more than 650,000 Creative Commons-licensed sound samples and effects. Search works on the website, and a clean JSON API is available with a free token. The natural stop when you need a clearly licensed sound clip or want to trace the origin of a widely reused sample. API · key
- Yandex Images Yandex's reverse-image search, regarded as best in class for faces and visually similar images, drawing on a multi-billion-image index. You submit an image URL and review matches in the web interface; there is no official API, and automated use trips captchas, so it works as a hand-off link. A standard OSINT stop when Google's reverse search comes up empty, especially for photographs of people. Images captcha
- Google Images / Lens The world's largest reverse-image and visual search, with Lens adding object and text recognition. It is web-only, there is no public API, so you upload an image or paste a URL yourself. Still the default first move for identifying an image's origin, its subject, or other places it appears online. Images captcha
- TinEye The oldest dedicated reverse-image engine, indexing more than 70 billion images and specializing in exact and edited copies, including the earliest indexed appearance of an image. The web search is free to use manually; the API is paid, so there is no automated route here. Especially valuable when you need to date an image or establish where it first surfaced. Images
- Bing Visual Search Microsoft's visual and reverse-image search. The web interface still works for matching an image against Bing's index, but use is manual only. It serves as a second opinion alongside Google and Yandex when reverse-searching an image. Degraded: the programmatic API was retired in August 2025 along with the rest of the Bing Search APIs. Images
- IQDB A combined reverse-image search across the major booru image boards. Danbooru, Gelbooru, Zerochan and others, covering tens of millions of tagged anime-style images. You give it an image or URL and it checks all the services at once; output is plain HTML with no API. The quickest way to find the catalogued, tagged source of an anime image when SauceNAO leaves doubt. Images
- ascii2d A Japanese reverse-image engine matching by color and feature signatures, particularly strong on artwork from Pixiv and Twitter. It is web-only, with no API. Worth a try for Japanese fan art when the mainstream reverse-search engines fail. Status is uncertain: it was unreachable from our infrastructure in testing, connections from outside Japan are often blocked, so expect access trouble. IP-blocked
- Karma Decay A reverse-image search scoped entirely to Reddit: give it an image and it surfaces prior submissions of the same picture across the site. It is a plain website with no API. Handy for tracing when and where an image first circulated on Reddit. Currently degraded, the site timed out in our checks, and its freshness is limited since Reddit locked down data access, so treat it as a long shot.
- PimEyes A face-recognition search engine covering roughly 3.5 billion face images from the open web: upload a photo of a face and it finds other pictures of the same person. It runs through the website only, and actually viewing results requires a paid subscription; there is no public API. Used in person-identification research when ordinary reverse-image search is not face-aware enough. captcha
- Photobucket A legacy image host holding billions of pictures from the earlier web, when it sat behind countless forum posts and blogs. There is no API, and much of the archive is hard to reach: hotlinked images often display placeholders and old albums are paywalled. Even so, it is occasionally the only surviving home of an old forum image, so a manual check can pay off. Degraded, expect broken or gated content frequently.
- Pinterest A visual-discovery service with billions of pinned images, which often preserve copies of pictures whose original sources have vanished. There is no public image-search API, the official one is OAuth-gated for approved business apps, and casual browsing hits login walls, so use it by direct link. Occasionally useful for finding a surviving copy of an image lost from its original site. IP-blocked
- Google Arts & Culture Google's portal to high-resolution artworks and cultural exhibits, offering millions of items from more than 2,000 museums and institutions. Everything runs through the web interface; there is no public API. Useful when researching artworks, artifacts or exhibitions, particularly when zoomable high-resolution imagery matters. Degraded only in the programmatic sense, its developer tooling was retired, so manual browsing is the route.
- Tube Archivarix Archivarix's own deleted-YouTube-video finder and a sibling service of this site. Given a video ID it hunts for live reuploads of the deleted video across mirror hosts, shows the video's archived metadata, and supports title search over its YouTube metadata index. Being first-party, it is fully integrated here with no key or setup needed. The first stop when a YouTube link is dead and you want either the video's details or a surviving copy. Video / audio / music
- Filmot A searchable index of YouTube metadata and subtitles spanning hundreds of millions of videos, including ones since deleted or made private. You can search video metadata and even words spoken in subtitles through the free web interface; the informal API needs a key issued by the operator. Its unique value is surfacing details of deleted YouTube videos, which often survive nowhere else. Degraded at present: video pages sit behind a Cloudflare gate against automated connections, so plan on browsing it manually. Video / audio / music captcha
- Ghostarchive (video) An on-demand, user-fed archiver notable for preserving YouTube and Twitter video as well as ordinary pages. There is no search API; you open archive pages directly by ID or browse its YouTube section. Worth checking in any deleted-video hunt alongside Filmot and the Wayback Machine, since it sometimes holds an actual playable copy rather than just metadata. Video / audio / music
- ANY.RUN An interactive online malware sandbox whose public report feed makes a large corpus of analyses freely browsable. You search and read public reports on the website; the API is limited to paid plans. Reach for it when you want to see how a suspicious file behaves when run, without executing it yourself.
- Hybrid Analysis The community portal of CrowdStrike's Falcon Sandbox, with a large public corpus of malware analyses searchable by file hash. The website itself is captcha-protected, so the workable route is its API, keyed via a free account. A solid stop for checking whether a hash has already been analyzed and what the sandbox observed. API · key captcha
- VirusShare An invite-only repository of tens of millions of malware samples, with hash lookup and sample download for members. Getting in means emailing the operator for an account; the API likewise needs the member key. Researchers turn to it when they need the actual sample behind a known hash and the open repositories do not have it. captcha
- BTDigg (btdig.com) A torrent search engine that crawls the BitTorrent DHT itself, indexing hundreds of millions of torrents and returning magnet links only. It is a plain HTML site with no API, promoted primarily as a Tor service. Good for finding content that is still seeded but listed on no conventional torrent index. Degraded: its clearnet address is flaky, so expect intermittent availability, and as ever with torrents, results are links only. Torrents IP-blocked
- SolidTorrents A DHT-based search engine over tens of millions of torrents and their metadata. It is web-only, with no official API, and its domain changes periodically. Another fallback for checking whether something still circulates on BitTorrent when the big indexes come up empty. Status is uncertain, it answered only via proxy in our checks, and the current live domain should be verified before relying on it. Torrents IP-blocked
- Snowfl A meta-search tool for torrents: rather than holding its own index, it queries multiple torrent indexers at once and aggregates the results. There is no API, so you use its web interface directly, from here it is surfaced as an outbound link only. Worth a stop when you want a single query spread across several torrent indexes instead of checking each one by hand, though it blocks some IP ranges and the site is heavily JavaScript-driven. IP-blocked
- TorrentGalaxy A general-purpose torrent index with an attached community, covering a large catalogue of torrents. It has no official API; you search it on its own website, which we link to rather than query. Researchers tracking files distributed via BitTorrent can use it as one of several indexes to check. Currently degraded: the site has faced instability and blocking pressure since 2025, so verify which domain is live before relying on it. IP-blocked
- 1337x One of the most-trafficked general torrent indexes on the web, with a very large catalogue. No official API exists, so searching happens on the site itself; we provide a link out only. It is a common first stop when looking for a torrent of almost any kind of file, but be aware it is ISP-blocked in many regions, so you may need a mirror or proxy to reach it. Torrents IP-blocked
- Bitmagnet (self-host) Unlike the other torrent entries, this is software you run yourself: a self-hosted crawler that listens to the BitTorrent DHT and builds your own torrent index. What it holds depends entirely on your own crawl, there is no public website or shared database to search. Once running, it exposes a local GraphQL and torznab interface on your own machine. Reach for it if you want an independent, self-controlled view of what circulates on the DHT rather than relying on public index sites.
- GH Archive A running record of public activity on GitHub: the public event timeline has been archived hourly since 2011, amounting to billions of events. It is a bulk dataset rather than a search site, you download gzipped JSON files per hour or query the public BigQuery dataset; there is no per-record search API or web search form. Researchers reach for it to analyze or reconstruct historical GitHub activity offline.
- Sourcegraph (public code search) Code search across millions of public repositories, letting you find where a snippet, identifier, or string appears anywhere in open-source code. You search through the web UI; a GraphQL endpoint also answers public searches without a token. Useful whenever a question spans many codebases at once rather than a single project. The future of the free public tier is uncertain after the company's closed-source pivot, so treat it as a convenience that may change.
- ccMixter A community site for Creative Commons-licensed music, remixes, samples, and a cappellas, on the order of tens of thousands of CC tracks. Searching and browsing happen on the site itself; a legacy query API exists but failed our checks. It is the place to look for reusable, openly licensed music stems and remixes. Status is currently uncertain: the site was unreachable in repeated probes and may be down, so check back if it doesn't load.
- GifCities (GeoCities GIFs) A dedicated search engine for animated GIFs salvaged from GeoCities, rebuilt in 2025 with semantic search. Type a term and it returns matching GIFs from the rescued corpus; its open search endpoint also lets this site query it directly. Reach for it when researching 1990s web aesthetics or hunting a specific piece of early-web imagery. Images API
- OoCities A mirror and resurrection of GeoCities pages saved before the service closed, browsable by the original neighborhood paths. There is no API; you browse or use the search form on the site itself, which we link to. Useful when you need to revisit a specific GeoCities homepage or wander the old neighborhood structure as it once was.
- GeoCities.ws Another GeoCities mirror, paired with free retro-style hosting, so archived pages sit alongside newly hosted ones. It has no API, use the browsing and search form on the site via the link provided. Handy as a second place to check when a GeoCities page is missing from other mirrors. Currently degraded, and it may present a CAPTCHA before letting you in. captcha
- ReoCities One of the original rescue mirrors created when GeoCities shut down in 2009, preserving pages saved during that effort. Browsing and searching happen on the site itself; there is no API, so we link to its search form. As one of several independent rescue projects, it may hold pages the other mirrors missed, which makes it worth checking when a page can't be found elsewhere.
- IA GeoCities Collection The Internet Archive's dedicated crawl of GeoCities, the same corpus that powers GifCities. It is reached through the Wayback Machine's CDX index scoped to geocities.com, so this site can query it directly and return lists of captures with timestamps. The most systematic way to find archived copies of a specific GeoCities URL, though queries can be slow. API
- Cameron's World Not a search tool but a web-art piece: a curated interactive collage assembled from GeoCities graphics and GIFs. You simply visit and scroll, there is no API or search function, so it is listed as a link. Worth a visit for a feel of GeoCities-era visual culture, or as a starting point when exploring what the amateur web of the 1990s looked like.
- ProtoWeb A service for vintage computing: an HTTP proxy that serves restored 1990s websites to period-appropriate old browsers. You point a vintage browser or emulator at the proxy and browse the early web from its cache of restored sites; there is no API and we link to it only. Of interest when you want to experience old sites on the software of their day. Currently degraded and may show a CAPTCHA. captcha
- oldweb.today Lets you open archived web pages inside emulated vintage browsers, all running in your modern browser, with content pulled from Wayback and other Memento archives. Use is entirely interactive on the site, pick a browser and a date, enter a URL; there is no API. Researchers reach for it when how a page rendered matters as much as what it said.
- Marginalia Search An independent search engine with its own crawler, focused on the small, old, text-heavy, non-commercial web. You can search on the site, and its open public API lets this site query it directly. Reach for it when hunting personal homepages, hobby sites, and long-form text from the web's quieter corners that mainstream engines tend to bury. Web full-text API
- Wiby A search engine devoted to hobbyist and early-web pages, complete with a 'surprise me' button that serves up a random old page. Search on the site itself, or via its simple JSON search endpoint, which this site can query directly. Good for serendipitous discovery of the classic web and for finding small hand-made pages on a topic. Web full-text API
- Mojeek An independent UK search engine that runs its own crawler rather than reselling Bing or Google results. You search through its normal web interface; a commercial paid API exists but is not used here, so we link out to the site. Useful when you want an index that doesn't simply mirror the big engines' view of the web.
- Stract An open-source search engine, funded by NLnet, designed to be customizable and self-hostable. Searching is done on its website; its public API is undocumented, so we link to the search form rather than query it. An option when you want an independent, transparent index, or your own instance of one.
- Million Short A search engine with a twist: it deliberately removes the most popular sites from the results so obscure pages can surface. You use its web search form directly, there is no API. Helpful for digging past the usual dominant domains when researching a topic. Currently degraded, and you may hit a CAPTCHA. captcha
- Teclis A non-commercial search index for the 'creative web' of small, independent sites, which now serves as Kagi's internal index. You search it through its own site; programmatic access exists only via Kagi's paid API, so it is linked here rather than queried. Try it when mainstream engines drown out the small sites you're actually after.
- Google Groups Usenet Google's archive of Usenet newsgroups, built on the DejaNews collection, with discussions reaching back to 1981. It has been frozen and read-only since 2024, and there is no API, you search and read through the Google Groups web interface we link to. Still one of the deepest places to find decades-old newsgroup threads and historical online discussion. Niche text
- Usenet Archives An independent, free web archive holding hundreds of millions of historical Usenet posts. Searching and reading happen on the site itself; there is no API, so it is presented here as a link. A useful alternative or complement to Google Groups when tracking down old newsgroup discussions.
- Narkive A long-running web archive of Usenet that lets you search and read historical newsgroup posts. There is no API; you use the search form on the site, which we link to. Worth checking alongside the other Usenet archives, since each covers and surfaces the material a little differently.
- textfiles.com Jason Scott's curated archive of the BBS era: text files, zines, phreaking documents, and BBS lists from the pre-web online world. It is a plain browsable site with no API, you explore it directly through the link provided. A core reference point for researching BBS culture and early online text.
- DiscMaster (files inside old discs) A standout tool that searches inside files: full-text and file-level search across millions of vintage files contained in old disc images stored on archive.org. It has a working search endpoint, so this site can query it directly as well as link you through. Reach for it when hunting a specific old shareware program, document, or image that shipped on a CD or floppy, content ordinary web search cannot see. Niche text API
- MobyGames A comprehensive catalogue of video game history, covering games across all platforms. Searching is done on the website; an API exists but requires a key. It is the standard reference when you need the catalogued record of a particular game. Currently degraded, the site may interpose a CAPTCHA. captcha
- IA MS-DOS / Software Library The Internet Archive's library of MS-DOS games and historical software, playable in-browser through emulation. It is searchable via the IA's advancedsearch interface, which this site queries directly, and items run right in your browser on archive.org. The quickest way to find and actually run a piece of vintage software without installing anything. Software / packages API
- My Abandonware A large database of abandonware, games from the 80s and 90s for DOS, Windows, Amiga, and consoles, with downloads alongside the catalogue entries. There is no API; you search and browse on the site itself via the link provided. A practical stop when you need to identify or obtain an old game that is no longer sold. Software / packages
- Europeana Europe's cultural heritage aggregator, bringing together more than 50 million digitized items from European galleries, libraries, archives, and museums. You search it on its website; an API is also available with a free key. The natural first stop when looking for digitized European cultural material across many institutions at once.
- DPLA The Digital Public Library of America aggregates digitized items from libraries, archives, and museums across the United States into one searchable index. Searching is done on its website; a free-key API is also available. Reach for it when looking for American historical materials without knowing which institution holds them.
- Smithsonian Open Access Open Access records from across all the Smithsonian museums, combining collection metadata with CC0 media that can be reused freely. Searching happens on the Smithsonian site; an API exists, keyed through api.data.gov. Valuable for finding reusable museum imagery and object records in one place. Currently degraded, expect possible CAPTCHA interruptions. captcha
- Met Museum API The Metropolitan Museum of Art's full collection records, with object metadata and open-access images. It offers a keyless public API, so collection data can be searched directly from this site without any account. Useful for art-historical research or for sourcing openly licensed images of museum objects. Probes currently rate it degraded, so responses may be inconsistent. Cultural heritage API
- Rijksmuseum API The Dutch national museum's collection online, with object metadata and high-resolution images. You search the collection on the museum's website; an API is available with a free key. A primary source when researching Dutch art and history or looking for high-quality images of the museum's objects.
- Chronicling America (LoC) The Library of Congress's collection of historic US newspaper pages from the 1700s through the 1960s, with OCR full text so the pages themselves are searchable. You search it through the loc.gov website, which can also return plain JSON without a key. Indispensable for finding period news coverage, names, and events in American newspapers. Note that the legacy chroniclingamerica.loc.gov address is deprecated, use loc.gov. API
- crt.sh (Cert Transparency) A queryable window into Certificate Transparency logs, the public record of TLS certificates, which makes it a way to uncover historical certificates and subdomains for a domain. You query it on the website, and the same query can return JSON without a key. Researchers use it to reconstruct a domain's infrastructure history. It is frequently overloaded and errors out under load, so retry if a query fails. Web pages / sites
- Censys Search An internet-wide scanning index of hosts and certificates, similar in role to Shodan. Searching is done through its web interface; an API with a free tier exists but requires a key, so it is linked here rather than queried. Relevant when investigating what a host or domain exposes to the internet and its certificate footprint. Currently degraded, the site may put a CAPTCHA in your way. captcha
- Shodan A search engine for internet-connected devices and services that also keeps historical service banners. You query it through its own web interface, which we link to; an official API exists but requires a key and is mostly paid. Researchers reach for it when tracing what services or software a host has exposed over time, as part of domain and infrastructure history work.
- Intelligence X Intelligence X is an OSINT search service covering pastes, darkweb content, data leaks and whois records, with a 'time machine' for material that has since been removed. You use it through its own website via our link; its API needs a key and the deeper data sits behind paid tiers. Worth a visit when you are chasing deleted pastes or leaked material that ordinary web archives never captured.
- Have I Been Pwned (pwned passwords) An index of accounts and credentials that have appeared in known data breaches. The Pwned-Passwords range endpoint is keyless, so our search can query it directly, while the fuller breach-lookup API is paid and individual account checks happen on the website itself. Use it to confirm whether a credential has surfaced in a leak. The service is flagged as degraded in our latest checks, so the website may be inconsistent to reach. API
- Ahmia (Tor index) A clearnet search engine that indexes Tor .onion hidden services, with abuse-filtered results. There is no API; you run queries on its own search form, which we link to. It is the practical way to discover onion-site content from the regular web, making it a starting point for darkweb-adjacent research.
- APKMirror An archive of historical Android APK versions, hosting signed and verified packages. There is no API, so you browse and search on the site itself via the link we provide. Reach for it when you need an older release of an Android app that current app stores no longer offer. Software / packages
- F-Droid The repository of free and open-source Android apps, which keeps a full version archive and builds packages reproducibly. Its package index is published as keyless JSON, so our search can query it directly as well as link you to the site. The place to look for current and historical versions of FOSS Android apps. Software / packages API
- PyPI The Python community's package index, holding every published version of each package along with structured metadata. You can search on the website, and its keyless JSON API lets our search pull package details directly. Useful for software archaeology: retrieving an old release of a library or checking a package's version history. Software / packages API
- npm registry The package registry behind the Node.js ecosystem, with full version history and downloadable tarballs for every package. The underlying registry API is keyless and our search queries it directly. Reach for it to trace a JavaScript package's release history or fetch an old version. Note that the npmjs.com website itself currently challenges traffic from datacenter networks, so deep links to package pages may show a verification step first. Software / packages API captcha
- OldMapsOnline A portal for discovering georeferenced historical maps held by institutions around the world, pointing you to the collection that holds each map. It offers a geographic (bounding-box) search API and IIIF endpoints, but we surface it as a link to its own search form. Currently flagged as degraded: the site may present a CAPTCHA before letting you in. Cultural heritage captcha
- David Rumsey Maps One of the premier digitized historical map collections, offering high-resolution scans, many of them georeferenced, on the Luna platform. Partial Luna and IIIF APIs exist, but the practical route is the site's own search interface, which we link to. A first stop when you need a high-quality scan of an old map.
- OSM / Overpass OpenStreetMap is the collaboratively edited map of the world; the Overpass API provides read-only structured queries over its data, and full history dumps preserve the map's past states. Overpass is keyless (with polite-use limits), and our search can query it directly. Relevant when you need to know what mapped features exist at a location, or to dig into the map's edit history. API
- GDELT A global news-monitoring project that tracks events and tone across world media, including television news. Its document API is keyless, so our search queries it directly in addition to linking to the project site. Use it to find news coverage of an event across many countries and outlets at once. API
- IA TV News Archive The Internet Archive's TV News collection makes US and international television news since 2009 searchable through closed-caption text. It is queried via the Internet Archive's keyless search API, which our site uses directly. Invaluable when you need to establish when and where something was said on television. Video / audio / music API
- Radio Garden An interactive globe for tuning into thousands of live radio stations around the world. An unofficial keyless API exists, which our search can query alongside linking to the globe interface itself. More a live listening tool than an archive, but handy for locating stations by place. API
- IA Live Music (etree) The Internet Archive's live music collection (etree) holds legal taper recordings of live concerts. It is searchable through the Archive's keyless search API, so results appear directly in our search as well as on archive.org. The place to look for recordings of a band's past shows. Video / audio / music API
- Zenodo CERN's general-purpose repository for research data and software, where every deposit receives a DOI. It has a keyless read API, but the service blocked our probing addresses, so for now we link to its own search instead of querying it. Reach for it when hunting datasets, software releases or supplementary materials behind published research. Its status is uncertain only from our vantage point; it is known to be operational elsewhere, so the link should work for you. Datasets
- Figshare A repository where researchers deposit datasets, figures, posters and other supplementary outputs. Its read API is keyless and our search queries it directly, or you can browse on the site. Useful when the data behind a paper was published separately from the paper itself. Datasets API
- Harvard Dataverse Harvard's instance of the Dataverse platform, a large repository of research data. The search API is keyless and our search queries it directly; the website offers the same search for browsing. A standard stop when looking for academic datasets. Datasets API
- OSF The Open Science Framework is a platform for research collaboration that hosts project materials, data and a preprint registry. Its read API is keyless, letting our search query it directly. Useful for finding preprints and research project materials that never appeared in journals. Datasets API
- Hugging Face Datasets The largest hub of machine-learning datasets (and models), each carrying rich metadata. A keyless API serves search results, which our site queries directly, and everything is browsable on the website. The default place to look when you need a published ML dataset or its documentation. Datasets API
- Archive of Our Own Archive of Our Own, run by the Organization for Transformative Works, is the largest fanfiction archive. It has no official API as a matter of policy, so we link you to its own search and filtering interface. Essential for any fan-works research. Currently flagged as degraded: expect a browser-verification step before the site loads. Niche text captcha
- FanFiction.Net A veteran fanfiction site whose collections date from the era before AO3. There is no API and the site sits behind Cloudflare, so it is presented as a link to its search form. Worth checking for older fan works from that earlier period. Flagged as degraded: you may need to pass a verification challenge to get in. captcha
- Fanlore A wiki run by the Organization for Transformative Works documenting fan history, fandoms and fan terminology. As a MediaWiki site it has a standard keyless API, but we currently surface it as a link to its own search. It is the reference work for understanding the fandom context around archived fan works. Flagged as degraded: the site may challenge your browser before loading. captcha
- WikiTeam dumps Preserved dumps of more than 600,000 wikis, including Fandom/Wikia communities and niche MediaWiki sites, stored on the Internet Archive. The collection is searchable via the Archive's keyless search API, which our site uses directly. The place to look when a wiki has vanished and you need its content back. Niche text API
- Discogs A crowdsourced database of music releases and physical pressings, paired with a marketplace. Its search API requires a token, so we surface it as a link to the site's own search. The standard reference for discographies and for identifying a specific pressing of a release. Flagged as degraded: a CAPTCHA may stand between you and the site. captcha
- MusicBrainz An open music encyclopedia that assigns canonical identifiers (MBIDs) to artists and releases. Its web-service API is keyless, rate-limited to one request per second, and our search queries it directly. The go-to source for clean, structured music metadata. Video / audio / music API
- Genius A database of song lyrics together with annotations. Its search API needs an OAuth token, so we link you to the site's own search instead of querying it. Useful for pinning down lyrics and the annotated context around them.
- CourtListener / RECAP A free archive of US case law, court documents gathered from PACER through the RECAP project, and oral-argument audio. Its REST search API works without a key (a token raises rate limits), and our search queries it directly. The first stop for tracking down American court records at no cost. Cultural heritage API
- PatentsView (USPTO) A USPTO patent data service whose distinguishing feature is disambiguated inventor and assignee records. The search API requires a free key and structured queries, so for now we link to its search interface rather than querying it directly. Useful when you want to follow a specific inventor's or company's patents rather than just keyword-match documents.
- Google Patents Full-text patent search spanning patent offices worldwide, plus non-patent literature. There is no official API (the data is separately available as a BigQuery public dataset), so you search on the site itself via our link. One of the most convenient single interfaces for global patent lookups. Academic papers
- Espacenet / EPO OPS The European Patent Office's worldwide patent search, accompanied by the Open Patent Services REST API. The API requires an OAuth key (a free tier exists), so we currently link to the Espacenet search interface instead. Use it to search the EPO's worldwide patent database. Flagged as degraded: the site may present a verification challenge before loading. captcha
- FamilySearch The world's largest free genealogy records archive, operated by the LDS Church. Its REST API requires an approved key plus login, so we link you to the site's own record search. The default free starting point for genealogical research.
- pouet.net (demoscene) The canonical database of the demoscene, documenting productions, parties and groups since the 1990s. Its data is published as periodic JSON dumps (data.pouet.net) that our search can draw on, alongside the site's own browsing and search. The reference source when researching demoscene releases or group histories. API
- 16colo.rs (ANSI/ASCII art) An archive of ANSI and ASCII artpacks from the BBS artscene, with material going back to the early 1990s. A keyless API is advertised, but our probe of it failed, so we link to the site's own browsing and search. The primary home of preserved BBS-era artwork; reach for it when researching the artscene or locating a specific artpack.
- ASCII Art Archive A large, categorized library of classic single-image ASCII art. There is no API; you browse the categories on the site itself via our link. Handy when you want examples of traditional ASCII artwork organized by subject.
- CyberLeninka (КиберЛенинка) The largest Russian-language open-access scholarly library: full-text journal articles across every discipline, free to read without registration. Search by title, author, or topic in the web interface, there is no official public API, so this site hands you a search link. The first stop for Russian-language papers that Western indexes (Crossref, OpenAlex) often miss. Academic papers
- НЭБ. National Electronic Library (rusneb.ru) Russia's National Electronic Library aggregates digitized books, periodicals, and dissertations from the Russian State Library and partner institutions. Public-domain scans can be read directly in the web viewer; in-copyright works are restricted to reading rooms. Search the collection on the website, useful for pre-revolutionary and Soviet-era Russian print that exists nowhere else online. The site was unreachable during our last check, so availability may vary.
- End of Term Web Archive A joint project of the Internet Archive, Library of Congress, and university partners that crawls US federal government websites at every presidential transition. It preserves .gov and .mil content that routinely disappears when administrations change, pages, reports, and datasets deleted from live sites. Browse the project portal to reach the term crawls, which replay through partner wayback instances. Reach for it when a government page vanished around an election.