Browse, filter, and download our complete collection of archived early-web content. Available in multiple formats for research, analysis, and preservation.
Full HTML source of all captured web pages from 1990 through 1999, including raw documents, server logs, and HTTP headers.
180,000 recovered GeoCities personal homepages with original layout, animations, and embedded media intact.
340,000 animated and static GIFs from the golden age of web animation — including separators, counters, buttons, and more.
Complete metadata index of all Web Rings discovered during crawling — including node links, themes, and member URLs.
Archived MIDI files, background music tracks, and embedded audio clips from personal homepages and commercial sites.
Over 2 million guestbook entries from personal homepages, fan sites, and community boards — raw text with timestamps.
Commercial and early e-commerce websites from the dot-com era, including Netscape-era layouts with frames and tables.
Pixel-perfect screenshots of 520,000 web pages rendered in their original browsers — Netscape, IE4, and Opera.
Identified technologies for every page: CMS, languages, frameworks, server types, and browser compatibility data.
Full HTML source of 892,000 pages with raw documents, server logs, and HTTP headers
180,000 recovered personal pages with original layouts and embedded media
340,000 animated and static GIFs — separators, counters, buttons, and web graphics
Complete metadata of 12,400 Web Rings with node links, themes, and member URLs
48,000 archived MIDI files, background tracks, and embedded audio clips
2.1 million guestbook entries from personal homepages and community boards
156,000 commercial sites from the dot-com era with Netscape-era layouts
520,000 pixel-perfect screenshots rendered in original browsers
Identified technologies for every page — CMS, languages, frameworks, server types
Download the entire archive or create a custom collection tailored to your research needs.
Query and download datasets programmatically through our RESTful API.
Our datasets are available in multiple formats including WARC, HTML tarballs, JSONL, Parquet (for structured data), PNG/TIFF (for screenshots), and FLAC (for audio). Each dataset page lists the exact format and includes schema documentation.
All archived content is provided under a Creative Commons BY-NC 4.0 license for non-commercial research. Commercial licensing is available — contact us for enterprise plans. All data includes cryptographic signatures for provenance verification.
The full archive can be downloaded via our batch download portal, torrent, or direct FTP. For the full 2.1TB dataset, we recommend our torrent option which distributes the load across our peer network. Torrent magnet link is available in the batch section above.
Yes! We welcome contributions of early-web content, personal homepages, and community pages. Upload through our submission portal and our team will verify, catalog, and add your content to the archive. Contributors receive a permanent attribution link.
We run crawls monthly and add new recovered content quarterly. The GeoCities recovery project has been our most active area — over 18,000 new pages are added each quarter as we discover and restore previously lost content.