What We Preserve & Why
The early web was fragile, ephemeral, and radically different from today's web. Our standards are designed to capture that era exactly as it existed, without modern interpretation or alteration.
🔒 Bit-Perfect Authenticity
We never modify source code, CSS, or assets. Tables, frames, inline styling, and deprecated tags are preserved exactly as authored. No auto-correction, no modernization.
🌐 Complete Context
Every archived page includes full metadata: capture timestamp, HTTP headers, referrer paths, server response codes, and linked asset trees to reconstruct the original environment.
📜 WARC Compliance
All captures follow the WARC 1.1 specification (ISO 28500), ensuring long-term interoperability with global digital preservation networks and academic repositories.
🛡️ Cryptographic Verification
Every archive file is signed with SHA-256 hashes and stored with immutable integrity checksums. Tamper detection alerts trigger automatic re-verification cycles.
Capture & Verification Pipeline
Our automated systems run 24/7 across distributed capture nodes. Each URL goes through a multi-stage verification process before entering the permanent archive.
archive::capture --url "http://example.geocities.com/90s"→ Fetch HTTP/1.0 response + headers→ Resolve all relative assets (IMG, AUDIO, EMBED, FRAMESET)→ Generate SHA-256 digest for each resource→ Package into WARC block with metadata envelope✓ VERIFIED & SEALED
Discovery
Crawlers follow sitemaps, web rings, guestbooks, and manual submissions to map link graphs.
Capture
Browser-neutral fetchers record raw HTML, binary assets, and server responses without execution.
Verification
Checksum validation, link integrity testing, and metadata cross-referencing ensure completeness.
Storage
Geo-redundant, cold-storage WARC archives with quarterly integrity audits and migration protocols.
Supported Era Technologies
We maintain legacy rendering profiles and parsing engines to support the full spectrum of late-80s through late-90s web standards.
| Category | Supported Formats | Status |
|---|---|---|
| Markup | HTML 1.0, 2.0, 3.2, 4.01 Transitional | FULL |
| Layout | Table layouts, FRAMESET, IFRAME, TARGET | FULL |
| Images | GIF (87a/89a), BMP, TIFF, early PNG | FULL |
| Audio | MIDI, AU, WAV, REALAUDIO, QUICKTIME | PARTIAL |
| Protocols | HTTP/1.0, FTP, Gopher, WAIS | FULL |
| Scripts | JavaScript 1.0-1.2, VBScript, CGI-BIN | SANDBOXED |
* Partial/Sandboxed formats are preserved in their raw form and rendered in isolated, period-accurate emulation environments to prevent security risks while maintaining visual fidelity.
Responsible Archiving
We operate under strict ethical guidelines to balance historical preservation with legal compliance and creator rights.
- • Copyright & DMCA: We honor valid takedown requests while preserving publicly accessible historical snapshots for fair use research. All archived content is clearly marked as historical reference material.
- • Privacy Protection: Personally identifiable information (PII) in guestbooks, email harvesters, and early forums is automatically redacted in public views, while maintaining integrity in restricted research copies.
- • Non-Commercial Policy: The archive is funded through institutional grants and donations. Content cannot be used for commercial resale, AI training datasets, or automated scraping without explicit licensing.
- • Transparent Operations: All capture logs, exclusion lists, and methodology updates are published quarterly. We welcome academic peer review and community oversight.
Submit or Request
Found a critical site fading from the web? Have a collection of floppy disks or CD-ROMs containing early web content? We accept digital submissions and coordinate physical media digitization.