🏛️ Wayback & CommonCrawl Recon

Query Wayback Machine & CommonCrawl archives to uncover forgotten endpoints, exposed configs, API keys, and shadow infrastructure.

Wayback checks run from your browser (quick) or your Archive Helper (full scans); CommonCrawl via the Max Intel proxy; nothing is stored · Last updated

PROGRESS
STATS
0
Total Found
0
After Dedup
0
High Value
0
Wayback
0
CommonCrawl
RESULTS

Discovered Endpoints

URL ↕ Type ↕ Status ↕ Last Seen ↕ Source ↕ Archive
SEO CONTENT

How do I find deleted pages and hidden endpoints on a website? Query archived URL indexes instead of the live site. The Internet Archive's Wayback Machine holds over 1 trillion archived pages (as of October 2025) and exposes them through its CDX API, while each monthly CommonCrawl index covers approximately 3–4 billion pages. Max Intel queries both, deduplicates the results, and marks URLs found in both sources as higher-confidence findings — surfacing forgotten endpoints, exposed config files, and API keys that were never meant to stay public.

SEARCH PANEL

How does archived endpoint reconnaissance uncover hidden attack surface?

Max Intel's Wayback & CommonCrawl Recon queries the Internet Archive's Wayback Machine CDX API and the CommonCrawl Index API to retrieve every URL ever crawled for a target domain. Web crawlers capture publicly accessible paths — including configuration files, admin panels, API documentation, and database backups that were briefly exposed before being removed. According to the OWASP Testing Guide v4.2, passive reconnaissance using web archives is a foundational step in application security testing because it reveals the historical attack surface that active scanning misses.

Why do archived URLs matter for security assessments?

A file removed from a live server may still be accessible at a different path, cached by a CDN, or retrievable from the archive itself. The SANS Institute penetration testing methodology recommends web archive enumeration as a standard reconnaissance technique because it reveals: endpoints that existed during development but were removed before production, configuration files that leaked credentials or internal hostnames, API versions that were deprecated but never fully decommissioned, and backup files that contain source code or database schemas. A 2024 HackerOne report found that approximately 18% of valid bug bounty submissions involved endpoints discovered through passive reconnaissance rather than active scanning — many of which were found in web archives.

What is the difference between Wayback Machine and CommonCrawl?

The Wayback Machine (operated by the Internet Archive since 2001) provides the largest single web archive, with over 1 trillion archived pages (as of October 2025). Its CDX API supports time-range filtering and returns HTTP status codes alongside URLs. CommonCrawl is a separate non-profit that publishes monthly web crawls as open datasets. Querying both sources provides broader coverage — CommonCrawl sometimes captures URLs that the Wayback Machine missed, and vice versa. Max Intel deduplicates results across both sources and marks URLs found in both as higher-confidence findings.

Wayback Machine CDX API
The programmatic interface to the Internet Archive's URL index, returning archived URLs with timestamps, HTTP status codes, and MIME types for any domain. Supports wildcard subdomains, date range filtering, and result collapsing by URL key.
CommonCrawl Index
A publicly accessible index of URLs captured during CommonCrawl's monthly web crawls. Each crawl index covers approximately 3–4 billion pages and can be queried via a REST API that returns URL, status, MIME type, and crawl metadata.
Attack Surface Mapping
The process of identifying all externally accessible endpoints, services, and data stores associated with an organization. Archived endpoint discovery contributes to attack surface mapping by revealing paths that no longer appear in DNS records, sitemaps, or live crawls.
High-Value Target
An archived URL classified as containing potentially sensitive information — such as .env files, database credentials, API documentation (Swagger/OpenAPI), version control metadata (.git/config), or server configuration files that may reveal internal architecture.
Related Tools

🏛️ Wayback & CommonCrawl Recon — Frequently Asked Questions

How do I search the Wayback Machine for OSINT?

This tool queries the Wayback Machine’s CDX index and CommonCrawl to surface archived URLs, deleted pages, and exposed endpoints a site has since removed — useful for recovering content that’s no longer live.

How does Wayback Machine and CommonCrawl reconnaissance work?

Max Intel runs in two modes. By default (quick check, nothing to install) it looks up a curated set of high-value paths — the site root, robots.txt, sitemap.xml, security.txt, and common admin, login, API, config and backup paths — one at a time through the Wayback Machine availability API, straight from your browser, and reports which ones have captures with dates and links. A quick "not found" is not proof a path was never archived. Enumerating every archived URL under the domain comes from the Wayback Machine CDX index and needs the free Max Intel Archive Helper, which runs those requests from your own browser and IP. CommonCrawl is a separate source, still queried through the Max Intel proxy in both modes. The tool deduplicates results across both sources, classifies each URL by type (config files, JS, admin panels, API endpoints, backups), and flags high-value targets like .env files, credentials, and database dumps.

What is the difference between a quick check and a full scan?

A quick check is the default and needs nothing installed: it runs about fifteen Wayback Machine availability lookups for the highest-value paths — root, robots.txt, sitemap.xml, security.txt, /admin, /login, /wp-login.php, /.env, /.git/config, /api, /backup, /old, /test, /staging and similar — and shows which already have captures, with dates and archive links. It is fast but partial: it never enumerates every archived URL, and a "not found" is inconclusive. A full scan enumerates every URL the Wayback Machine CDX index holds under the whole domain and its subdomains, which needs the free Max Intel Archive Helper so the requests run from your own browser and IP instead of a shared proxy. Install the helper when you need the complete historical URL list rather than a spot check.

What types of sensitive files can archived endpoint discovery find?

Common high-value findings include exposed .env files with API keys, wp-config.php with database credentials, .git/config revealing repository structure, Swagger/OpenAPI documentation exposing internal APIs, database backups (.sql, .sql.gz), server configuration files (php.ini, web.config), and admin panel login pages. Even if these files have since been removed from the live site, the archived URLs confirm they once existed.

Is Wayback Machine OSINT reconnaissance legal?

Querying the Wayback Machine and CommonCrawl APIs for publicly archived URLs is legal — these are public data sources that archive the open web. However, using discovered endpoints to access live systems without authorization would violate computer fraud laws. This tool is intended for authorized security testing, bug bounty programs, and attack surface assessment of domains you own or have permission to test.

Wayback Delta Analyzer

Last updated:

Fetches archived versions of the current page from the Wayback Machine CDX API and diffs them against the live version. Highlights removed content, deleted links, changed contact information, edited paragraphs, and modified metadata. Free alternative to ChangeTower and Visualping ($15-100/mo).

Drag to your bookmarks bar:

Diff with Wayback
1
Install — drag to bookmarks bar
2
Visit any webpage
3
Click — fetches Wayback snapshots and compares against the live page

Runs on any website — all processing in your browser.

📜

Install the bookmarklet, then use it on any website

Wayback Delta Analyzer

Websites change constantly — content is removed, contacts disappear, links break. The Wayback Delta Analyzer captures the current state of any page and compares it against the most recent Wayback Machine archive, surfacing everything that changed.

OSINT Applications

Deleted contact information, removed employee pages, changed company descriptions, and broken links all provide intelligence. Content removal often indicates something the organization wants to hide.

CDX API
The Wayback Machine's index API that returns a list of all archived snapshots for a given URL.

📜 Wayback Delta Analyzer — FAQ

How far back does it compare?

It fetches up to 20 snapshots from the Wayback Machine and compares the live page against the most recent archive. The timeline shows all available snapshots.

Why are some links shown as removed/added?

Links are compared by exact URL. Minor changes (trailing slashes, parameter order) will appear as removals/additions even if the destination is the same.

Does it work on JavaScript-rendered pages?

The live page capture reads the rendered DOM. The archived version depends on what the Wayback Machine captured — it may not include JS-rendered content.

Can I compare two specific archive dates?

Currently it compares live vs most recent archive. Multi-snapshot comparison is planned for a future update.

What if the page has never been archived?

The tool will report no snapshots found. Consider triggering an archive.today snapshot first for future comparisons.

How this tool fetches data: most lookups go straight from your browser to the public source named in the results. Where a source blocks browser requests (cross-origin restrictions) or the direct call fails, the value you entered is passed through a proxy operated by Max Intel and forwarded to that source instead. Either way the request leaves your browser; the proxy stores nothing against you and no account is involved. Full detail in the privacy policy →