Download & Access Archive Data

Bulk exports, raw datasets, API endpoints, and verified indexes. Access 4.2M+ preserved pages from the dawn of the internet in researcher-ready formats.

Early Web Full Corpus
1990–1999 • Complete Archive
WARC HTML 842 GB

Comprehensive dump of all verified pages, assets, and metadata from the first decade of the WWW. Includes original timestamps and HTTP headers.

GeoCities Recovered Pages
1994–2009 • 182,400 Pages
HTML JSON 48 GB

Reconstructed personal homepages, community themes, and web rings. Structured JSON metadata included for each neighborhood and theme.

Sample Archive Pack
Multi-era • 1,200 Pages
HTML 240 MB CSV

Curated sample set spanning 1992-1998. Ideal for testing parsers, building demos, or academic exploration without heavy bandwidth usage.

Dot-Com Business Pages
1995–2001 • Corporate & Startup
WARC 156 GB HTML

Preserved corporate websites, early e-commerce layouts, and Y2K-era marketing pages. Includes navigation structures and early CSS experiments.

Raw WARC Streams
Continuous • Unprocessed
WARC 2.1 TB

Unfiltered capture streams from original crawlers. Best for advanced archival reconstruction and protocol analysis.

WebRing Networks
1996–2001 • 4,200 Sites
JSON HTML 8.4 GB

Graph-structured dataset mapping WebRing connections, adjacency lists, and thematic clusters across early community networks.

Public API Endpoints
RESTful • JSON/CSV
JSON Streaming

Query archived pages by URL, date range, or metadata. Includes pagination, rate limits, and batch export capabilities.

Choose Your Access Level

Structured tiers designed for hobbyists, academic researchers, and commercial enterprises.

Public
Free / forever

For students, hobbyists, and casual explorers.

  • Sample datasets (≤500 MB)
  • Basic search & preview
  • Standard API (100 req/day)
  • CC0 licensed content
Enterprise
Custom

For AI training, commercial archives, and partners.

  • Raw WARC streams
  • Dedicated endpoints
  • Unlimited API access
  • Commercial licensing & SLA

Data Formats & Integrity

Parameter Specification Details
Primary Format WARC 1.1 Web Archive standard with full HTTP transaction preservation
Metadata JSON-LD / CSV Semantic markup, timestamps, crawl paths, and deduplication flags
Checksums SHA-256 Every file verified. .sha256sum manifests included per dataset
Compression Zstandard / GZIP Optimized for archival storage and fast decompression
API Protocol REST + GraphQL See API Documentation for rate limits & schemas
CLI Tool 1990-archive-cli npm i -g 1990-archive-cli • Authenticated bulk downloads

Usage Guidelines & Licensing

All archived content is preserved under a CC0 1.0 Universal dedication where copyright has expired or explicitly waived. Commercial usage of verified datasets requires attribution to "1990 Web Archive" and compliance with our Data Policy.

Academic researchers must cite the archive using the provided persistent identifiers (DOIs). Automated scraping outside of approved API endpoints is prohibited and may result in access revocation.

CC0 1.0 Academic Citation Required No Unauthorized Scraping Commercial License Available
"} ```