4.2M
Total Pages
892K
HTML Documents
340K
GIFs Recovered
180K
GeoCities Pages
2.1TB
Total Size
34K
Active Researchers
📄 HTML 48.2 GB

Complete 1990-1999 HTML Corpus

Full HTML source of all captured web pages from 1990 through 1999, including raw documents, server logs, and HTTP headers.

📅 1990–1999 📄 892,000 docs 🏷️ v3.2
raw-html http-headers server-logs
📄 HTML 12.8 GB

GeoCities Personal Homepages

180,000 recovered GeoCities personal homepages with original layout, animations, and embedded media intact.

📅 1994–2009 📄 180,000 pages 🏷️ v2.1
geocities personal animated-gif
🖼️ Images 67.4 GB

Retro GIF Collection (1990–1999)

340,000 animated and static GIFs from the golden age of web animation — including separators, counters, buttons, and more.

📅 1990–1999 🖼️ 340,000 files 🏷️ v4.0
animated-gif web-buttons counters
8.6 GB

Web Ring Directory Index

Complete metadata index of all Web Rings discovered during crawling — including node links, themes, and member URLs.

📅 1995–2000 🔗 12,400 rings 🏷️ v1.8
web-ring metadata graph-data
🎵 Media 15.3 GB

Web MIDI & Audio Collection

Archived MIDI files, background music tracks, and embedded audio clips from personal homepages and commercial sites.

📅 1993–1999 🎵 48,000 files 🏷️ v2.3
midi background-audio wav
📝 Text 3.2 GB

Guestbook Entries Corpus

Over 2 million guestbook entries from personal homepages, fan sites, and community boards — raw text with timestamps.

📅 1994–2001 📝 2.1M entries 🏷️ v1.5
guestbook user-generated natural-language
📄 HTML 22.1 GB

Dot-com Boom Era Pages (1995–2001)

Commercial and early e-commerce websites from the dot-com era, including Netscape-era layouts with frames and tables.

📅 1995–2001 📄 156,000 pages 🏷️ v3.0
dot-com ecommerce frames
🖼️ Images 34.7 GB

Early Web Screenshot Archive

Pixel-perfect screenshots of 520,000 web pages rendered in their original browsers — Netscape, IE4, and Opera.

📅 1993–1999 📸 520,000 renders 🏷️ v2.7
screenshots rendered netscape
5.1 GB

Technology Stack Fingerprint Database

Identified technologies for every page: CMS, languages, frameworks, server types, and browser compatibility data.

📅 1990–1999 🔧 892K entries 🏷️ v1.2
tech-stack fingerprint classification
📄

Complete 1990-1999 HTML Corpus

Full HTML source of 892,000 pages with raw documents, server logs, and HTTP headers

48.2 GB GZIPPED TAR
📄

GeoCities Personal Homepages

180,000 recovered personal pages with original layouts and embedded media

12.8 GB JSONL + ASSETS
🖼️

Retro GIF Collection (1990-1999)

340,000 animated and static GIFs — separators, counters, buttons, and web graphics

67.4 GB TAR.GZ
🏷️

Web Ring Directory Index

Complete metadata of 12,400 Web Rings with node links, themes, and member URLs

8.6 GB CSV + GRAPHML
🎵

Web MIDI & Audio Collection

48,000 archived MIDI files, background tracks, and embedded audio clips

15.3 GB ZIP + FLAC
📝

Guestbook Entries Corpus

2.1 million guestbook entries from personal homepages and community boards

3.2 GB PARQUET
📄

Dot-com Boom Era Pages (1995-2001)

156,000 commercial sites from the dot-com era with Netscape-era layouts

22.1 GB WARC + HTML
📸

Early Web Screenshot Archive

520,000 pixel-perfect screenshots rendered in original browsers

34.7 GB PNG + TIFF
🔧

Technology Stack Fingerprint Database

Identified technologies for every page — CMS, languages, frameworks, server types

5.1 GB JSON + SQLite

Batch Download

Download the entire archive or create a custom collection tailored to your research needs.

🌐

Full Archive

All 4.2M pages, 2.1TB total

📅

By Decade

1990s or 2000s subsets

🏷️

By Category

HTML, images, media, text

⚙️

Custom

Build your own dataset

Programmatic Access via API

Query and download datasets programmatically through our RESTful API.

bash — search and download
$ # Search the archive for 1995 GeoCities pages $ curl https://api.1990webarchive.com/v1/datasets \ -d "{'era': '1990s', 'site': 'geocities', 'format': 'html'}" \ -H "Authorization: Bearer YOUR_API_KEY" $ # Returns: { "status": "success", "total_matches": "18,432", "dataset_size": "12.8 GB", "download_url": "https://dl.1990webarchive.com/datasets/geocities-1990s-v3.tar.gz", "expires_at": "2025-01-15T23:59:59Z" }
GET
/v1/datasets
List all available datasets with filters for era, type, and size range.
GET
/v1/datasets/{id}
Get detailed metadata for a specific dataset including checksums and schema.
POST
/v1/datasets/search
Query the archive with custom filters and receive a downloadable dataset link.
GET
/v1/datasets/{id}/stream
Stream dataset contents line-by-line for real-time processing pipelines.
DELETE
/v1/datasets/{id}
Delete a previously downloaded dataset from our CDN (after 24h cooldown).

Frequently Asked Questions

Our datasets are available in multiple formats including WARC, HTML tarballs, JSONL, Parquet (for structured data), PNG/TIFF (for screenshots), and FLAC (for audio). Each dataset page lists the exact format and includes schema documentation.

All archived content is provided under a Creative Commons BY-NC 4.0 license for non-commercial research. Commercial licensing is available — contact us for enterprise plans. All data includes cryptographic signatures for provenance verification.

The full archive can be downloaded via our batch download portal, torrent, or direct FTP. For the full 2.1TB dataset, we recommend our torrent option which distributes the load across our peer network. Torrent magnet link is available in the batch section above.

Yes! We welcome contributions of early-web content, personal homepages, and community pages. Upload through our submission portal and our team will verify, catalog, and add your content to the archive. Contributors receive a permanent attribution link.

We run crawls monthly and add new recovered content quarterly. The GeoCities recovery project has been our most active area — over 18,000 new pages are added each quarter as we discover and restore previously lost content.