๐Ÿ•ท๏ธ Crawler Tool

Deep-crawl and archive web pages from the early internet era

url
geocities.com angelfire.com tripod.com yahoo.com (1995) altavista.com amazon.com (1995)
๐ŸŒ

First Web Page

CERN's original 1990 page

๐Ÿ“š

GeoCities Recovery

Mass crawl of GeoCities pages

๐Ÿ”—

WebRing Archive

Crawl a WebRing chain

๐Ÿ“„

W3C Specifications

HTML 2.0 / 3.2 specs

๐ŸŽฏ Era Detection (Optional)

๐Ÿ“Ÿ
Pre-Web
1989โ€“1993
๐Ÿ’พ
Web 1.0 Early
1994โ€“1996
๐ŸŒˆ
GeoCities Era
1997โ€“1999
๐Ÿ’ฟ
Dot-Com Boom
2000โ€“2001
๐Ÿ“ฑ
Post Crash
2002โ€“2004
๐Ÿ”ฎ
Any Era
No filter

โš™ Crawl Settings

โ–พ
Crawl Depth Max link depth to follow
Max Pages Cap on total pages to crawl
Politeness Delay Seconds between requests
User Agent Browser identity for requests

๐Ÿ“„ Content Filters

โ–พ
Content Types
โ˜‘ HTML
โ˜‘ GIF
โ˜‘ JPEG
โ˜ MIDI
โ˜ AVI
โ˜ PDF
โ˜ CSS
โ˜ JS
Follow Redirects
Respect robots.txt
Archive Subdomains
Detect Era Metadata

โš™ Crawl in Progress

0
Pages
0 KB
Size
0
Errors
0s
Elapsed
0% complete Est. remaining: calculating...
Recent Crawled URLs
crawler-log โ€” 1990webarchive
[--:--:--] INFO Crawler initialized. Ready to accept crawl jobs.
[--:--:--] INFO Configure your settings above and enter a URL to begin archiving.
[--:--:--] INFO Supports: HTML, GIF, JPEG, MIDI, AVI, and legacy formats.
[--:--:--] INFO Tip: Use batch mode to crawl multiple URLs at once. ๐Ÿ“ฅ

๐Ÿ“Š Crawl Results 0

URL Status Format Size Detected Era Crawled At Actions
๐Ÿ•ธ๏ธ

No crawl results yet

Enter a URL above and start a crawl to see archived pages appear here. Each page is preserved with its original format and metadata.