What We Preserve & Why

The early web was fragile, ephemeral, and radically different from today's web. Our standards are designed to capture that era exactly as it existed, without modern interpretation or alteration.

🔒 Bit-Perfect Authenticity

We never modify source code, CSS, or assets. Tables, frames, inline styling, and deprecated tags are preserved exactly as authored. No auto-correction, no modernization.

🌐 Complete Context

Every archived page includes full metadata: capture timestamp, HTTP headers, referrer paths, server response codes, and linked asset trees to reconstruct the original environment.

📜 WARC Compliance

All captures follow the WARC 1.1 specification (ISO 28500), ensuring long-term interoperability with global digital preservation networks and academic repositories.

🛡️ Cryptographic Verification

Every archive file is signed with SHA-256 hashes and stored with immutable integrity checksums. Tamper detection alerts trigger automatic re-verification cycles.

Capture & Verification Pipeline

Our automated systems run 24/7 across distributed capture nodes. Each URL goes through a multi-stage verification process before entering the permanent archive.

# Standard capture workflow
archive::capture --url "http://example.geocities.com/90s"
Fetch HTTP/1.0 response + headers
Resolve all relative assets (IMG, AUDIO, EMBED, FRAMESET)
Generate SHA-256 digest for each resource
Package into WARC block with metadata envelope
VERIFIED & SEALED
01

Discovery

Crawlers follow sitemaps, web rings, guestbooks, and manual submissions to map link graphs.

02

Capture

Browser-neutral fetchers record raw HTML, binary assets, and server responses without execution.

03

Verification

Checksum validation, link integrity testing, and metadata cross-referencing ensure completeness.

04

Storage

Geo-redundant, cold-storage WARC archives with quarterly integrity audits and migration protocols.

Supported Era Technologies

We maintain legacy rendering profiles and parsing engines to support the full spectrum of late-80s through late-90s web standards.

Category Supported Formats Status
Markup HTML 1.0, 2.0, 3.2, 4.01 Transitional FULL
Layout Table layouts, FRAMESET, IFRAME, TARGET FULL
Images GIF (87a/89a), BMP, TIFF, early PNG FULL
Audio MIDI, AU, WAV, REALAUDIO, QUICKTIME PARTIAL
Protocols HTTP/1.0, FTP, Gopher, WAIS FULL
Scripts JavaScript 1.0-1.2, VBScript, CGI-BIN SANDBOXED

* Partial/Sandboxed formats are preserved in their raw form and rendered in isolated, period-accurate emulation environments to prevent security risks while maintaining visual fidelity.

Responsible Archiving

We operate under strict ethical guidelines to balance historical preservation with legal compliance and creator rights.

  • Copyright & DMCA: We honor valid takedown requests while preserving publicly accessible historical snapshots for fair use research. All archived content is clearly marked as historical reference material.
  • Privacy Protection: Personally identifiable information (PII) in guestbooks, email harvesters, and early forums is automatically redacted in public views, while maintaining integrity in restricted research copies.
  • Non-Commercial Policy: The archive is funded through institutional grants and donations. Content cannot be used for commercial resale, AI training datasets, or automated scraping without explicit licensing.
  • Transparent Operations: All capture logs, exclusion lists, and methodology updates are published quarterly. We welcome academic peer review and community oversight.

Submit or Request

Found a critical site fading from the web? Have a collection of floppy disks or CD-ROMs containing early web content? We accept digital submissions and coordinate physical media digitization.

Submit URL for Capture Research Access Request
"}