text
· 73 Archive · 11
Cluster
Books 10
- Open Library The Internet Archive's open catalog of books, a page per book, edition and author with covers and metadata, plus digital lending where available. Open
- Internet Archive Texts The Internet Archive's library of 40M+ scanned books, magazines and documents, many readable in full, some through controlled digital lending. Open
- HathiTrust A shared university digital library of 18M+ scanned volumes, full text where copyright allows, searchable catalog metadata for the rest. may require captcha Open
- Wikisource Wikimedia's free library of source texts, public-domain books, historical documents, speeches and laws, transcribed and proofread by volunteers. Open
- Google Books API Google's huge book index, metadata, covers and searchable previews for millions of titles; broad coverage by ISBN, title or author. Open
- Project Gutenberg (Gutendex) Project Gutenberg, ~75,000 free public-domain ebooks (classic literature and more) you can download in several formats, no account needed. Open
- Standard Ebooks Carefully hand-made, well-typeset editions of public-domain books, free, modern and proofread, unlike rough auto-generated scans. may be IP-blocked Open
- DOAB (Open Access Books) The Directory of Open Access Books, peer-reviewed scholarly books that are free to read and download, across every discipline. Open
- Anna's Archive A search across the major shadow libraries at once (Library Genesis, Sci-Hub, Z-Library and more), one place to find books and papers. Links to copyrighted works; opens externally. may require captcha Open
- Library Genesis (LibGen) Library Genesis, a long-running shadow library of books, textbooks and articles. Links to copyrighted material; opens externally. may be IP-blocked Open
Social 3
- Wayback, reddit.com Wayback Machine snapshots of Reddit threads and profiles, independent of Reddit's API, so they survive deletions and blackouts. Open
- Desuarchive (4chan archive) A searchable archive of several 4chan boards (/a/, /int/, /k/, /tg/, /wsg/…), keeping threads and images long after 4chan purges them; JSON API. Open
- 4plebs (4chan archive) A long-running searchable 4chan archive for boards like /pol/, /x/, /tv/ and /adv/, preserves threads and images after they're deleted; JSON API. may be IP-blocked Open
Software / packages 7
- Software Heritage A non-profit archive preserving all public source code ever written, the 'Library of Alexandria of code', with permanent IDs for files and commits. Open
- IA MS-DOS / Software Library Thousands of classic MS-DOS games and historical software you can play right in the browser, hosted by the Internet Archive. Open
- npm registry The registry behind Node.js / JavaScript packages, full version history and downloadable tarballs for any package. may require captcha Open
- PyPI The official Python package repository, every version of every library, with JSON metadata, so you can fetch or inspect any past release. Open
- F-Droid An app store for free and open-source Android software that keeps every past version, with reproducible, verifiable builds. Open
- APKMirror An archive of older Android app versions, signed, verified APKs so you can install or roll back to a specific release. Open
- My Abandonware A large database of 'abandonware', old 1980s-90s DOS, Windows, Amiga and console games, with details and downloads. Open
Web full-text 3
- Wiby A search engine for the old-style personal and hobbyist web, simple, handmade pages like the early 2000s; a 'surprise me' button serves a random one. Open
- Yandex (cache / SERP) Yandex search, one of the last major engines that still serves a cached copy of a page from its results, handy when the live page is down or has changed. may require captcha Open
- Marginalia Search An independent search engine that deliberately surfaces the small, old, text-first, non-commercial web, the opposite of SEO-tuned results. Open
Video / audio / music 13
- Tube Archivarix Archivarix's own tool for deleted YouTube videos, finds live reuploads across mirror hosts and shows the video's archived details. (A sibling of this site.) Open
- Filmot A search engine over YouTube metadata and subtitles, find videos, even deleted or private ones, by words spoken or in the description. may require captcha Open
- Internet Archive (Video) The Internet Archive's huge open video collection, old films, TV, ads and ephemera, free to watch, with a keyless search and metadata API. Open
- Dailymotion A major video-sharing site with an open, key-free API and reliable search, in our testing the most programmatically searchable place to find reuploaded videos. Open
- Odysee A video platform on the LBRY network, openly accessible with no captcha, a useful alternative host for videos removed elsewhere. Open
- Rumble A US video-sharing site often used for content removed from YouTube; it blocks data-center IPs, so you may need to open it yourself. may be IP-blocked Open
- BitChute An independent video host frequently used for content taken down elsewhere; its search is captcha-gated. may require captcha Open
- Bilibili China's largest video-sharing site, an enormous catalog, but search is captcha- and region-gated from outside China. may require captcha Open
- Internet Archive (Audio) / Live Music The Internet Archive's open audio collection, including the Live Music Archive of legally tradeable concert recordings. Open
- IA TV News Archive A searchable archive of US and international TV news since 2009, find clips by what was said, using the closed-caption text (Internet Archive). Open
- IA Live Music (etree) The Live Music Archive, tens of thousands of legally tradeable concert recordings from bands that allow taping (Internet Archive). Open
- MusicBrainz An open, community-built music encyclopedia, the canonical IDs and metadata for artists, albums and tracks (the 'Wikipedia of music'). Open
- RuTube Russia's main video-sharing platform, worth checking for reuploads of deleted videos, especially Russian-language content. may require captcha Open
Academic papers 15
- Crossref REST API The official registry behind journal-article DOIs, authoritative metadata (title, authors, journal, references) for ~165M scholarly works. Look up a DOI or a title. Open
- DataCite REST API The DOI registry for research outputs beyond papers, datasets, software and preprints, with their citation metadata. Open
- dblp (CS bibliography) The definitive computer-science bibliography, a curated, high-quality index of CS papers, authors and conferences (~7M records). Open
- arXiv (export API) The open preprint server for physics, math, CS, biology and more (~2.5M papers), free full-text PDFs, often months ahead of formal publication. Open
- bioRxiv The free preprint server for biology, run by Cold Spring Harbor Laboratory, full-text research papers shared before peer review. may require captcha Open
- medRxiv The free preprint server for medical and health research (sister site to bioRxiv), clinical and public-health studies shared ahead of peer review. may require captcha Open
- Europe PMC Europe's open database of life-sciences and biomedical literature (EMBL-EBI), abstracts plus millions of free full-text articles. Open
- DOAJ The Directory of Open Access Journals, a vetted index of free, peer-reviewed journals and their articles, so you reach the real paper without a paywall. Open
- OpenAIRE Graph The EU's open-science knowledge graph, links research papers, datasets, software and funding across Europe's repositories. Open
- OpenAlex A free, open index of the world's research, ~480M papers, authors, institutions and journals, linked by citations. The open successor to Microsoft Academic. Open
- Semantic Scholar (Academic Graph) Allen AI's free research search and citation graph (~220M papers), finds related work and shows the sentence-level context in which a paper is cited. Open
- PubMed / PMC (NCBI E-utilities) The U.S. National Library of Medicine's index of biomedical literature (~37M citations), linked to free full text in PubMed Central. Open
- scholar.archive.org / fatcat Internet Archive Scholar, full-text search across 25M+ open-access research papers, including ones rescued from journals that have since vanished. Open
- Google Patents Search patents from offices worldwide plus related technical literature, full text and PDFs, free. Open
- CyberLeninka (КиберЛенинка) A Russian open-access scholarly library, full text of peer-reviewed journal articles, free to read (mostly Russian-language research). Open
Cultural heritage 3
- Met Museum API The Metropolitan Museum of Art's open collection, metadata and high-resolution, free-to-use images for hundreds of thousands of artworks. Open
- CourtListener / RECAP Free US case law, court filings (via RECAP/PACER) and oral-argument audio, a non-profit legal-research archive. Open
- OldMapsOnline A search portal for historical maps held by libraries worldwide, find old, georeferenced maps of a place by location. may require captcha Open
Images 3
- Wikimedia Commons Wikipedia's media library, 100M+ freely licensed images, sounds and videos anyone can reuse; full search and API. Open
- Openverse A search engine for openly licensed media, 800M+ Creative-Commons and public-domain images and audio you can legally reuse (by Automattic/WordPress). Open
- GifCities (GeoCities GIFs) A search engine for the animated GIFs rescued from 1990s GeoCities home pages, a slice of early-web culture, by the Internet Archive. Open
Niche text 4
- DiscMaster (files inside old discs) Full-text and file search inside millions of vintage files, software, documents, images, extracted from old CD/disk images on the Internet Archive. Open
- WikiTeam dumps Preserved backups of 600k+ wikis. Fandom/Wikia communities and niche MediaWikis, saved to the Internet Archive before they disappear. Open
- Google Groups Usenet Google's archive of Usenet newsgroup discussions back to 1981 (the old DejaNews corpus), read-only and frozen since 2024. Open
- Archive of Our Own Archive of Our Own, the largest fan-fiction archive, run by a non-profit; browse or search millions of fan works (no official API). may require captcha Open
Datasets 5
- Harvard Dataverse Harvard's installation of Dataverse, a large open repository for sharing and citing research datasets. Open
- Figshare An open repository for research outputs that don't fit a paper, datasets, figures, posters and supplementary files, each citable. Open
- OSF The Open Science Framework, where researchers share projects, data, materials and preprints, and register study plans. Open
- Zenodo CERN's open repository for any research output, datasets, software, papers, posters, each given a citable DOI. Open
- Hugging Face Datasets Hugging Face, the largest hub of machine-learning datasets (and models), with searchable, well-documented metadata. Open
Torrents 7
- The Pirate Bay (apibay) The longest-running general torrent index; its apibay.org backend returns results as plain JSON. Links to copyrighted material; opens externally. Open
- Nyaa The main torrent index for anime and other East-Asian media (often with subtitles); its RSS feeds double as a simple API. Open
- 1337x One of the most popular general torrent indexes. Links to copyrighted material; may be IP-blocked, and opens externally. may be IP-blocked Open
- Torrents-CSV A simple, open torrent search built on a community-maintained dataset, no ads or tracking, with a clean JSON API. Open
- Academic Torrents A torrent service for sharing large research datasets and academic data, a legitimate way to distribute multi-gigabyte scientific files. Open
- SolidTorrents A torrent search engine built from DHT/metadata crawling, aggregates magnet links from across the network. may be IP-blocked Open
- BTDigg (btdig.com) A torrent search engine that crawls the BitTorrent DHT network directly, returns magnet links only, with no site hosting the files. may be IP-blocked Open