CrawlLedger
Append-only observatory · live pilot

A public change log for the machine-readable web.

CrawlLedger preserves what domains publish in robots.txt, Content Signals, llms.txt, and RSL licensing files—plus when it was observed, how it parsed, and how each record links to the one before it.

20watched domains
456observations
61unique artifacts
2parsed changes

Evidence that stays inspectable

A normal checker tells you what a file says now. CrawlLedger keeps the observation metadata needed to distinguish a current result, an unchanged response, a parse failure, and a fetch skipped by robots policy.

01 / Observe narrowly

One queued domain at a time, limited to public policy files in a reviewed 20-domain pilot.

02 / Preserve provenance

Content hashes, bounded response metadata, parser versions, and per-domain signal chains remain attached to every event.

03 / Report carefully

Pages use factual states such as observed, not observed, unverified, parse failed, or skipped—not legal conclusions.

One robots.txt file, many different AI systems

Training crawlers, search crawlers, user-triggered fetchers, and control tokens are not interchangeable. CrawlLedger keeps their first-party descriptions and observed policy states separate.

Content use

Understand the emerging Content-Signal fields for search, AI input, training, and reuse.

Recent parsed artifact changes

Only post-baseline byte changes that produced a parsed result are shown here. Changing error pages and routine 304 responses are excluded.

usatoday.com

robots.txt · Observed and parsed

9bf8ba4edf17…artifact hash
usatoday.com

robots.txt · Observed and parsed

b0d291be5914…artifact hash

Recently observed domains

Each domain page is a canonical, crawlable landing page for the public metadata already available from the JSON API.

anthropic.com

2 signal types · last observed Aug 15, 2026

24records
openai.com

2 signal types · last observed Aug 15, 2026

24records
cloudflare.com

2 signal types · last observed Aug 15, 2026

24records
rslcollective.org

3 signal types · last observed Aug 15, 2026

48records
rslstandard.org

3 signal types · last observed Aug 15, 2026

39records

Built for publishers, researchers, and agents

The archive is useful whether you need to audit a crawler rule, cite a historical observation, monitor emerging licensing signals, or consume structured metadata from an API.

Free policy checker

Paste a robots.txt file and evaluate the tracked AI crawler identifiers without triggering a network request or storing the input.

Open the checker →

Source-backed guides

Compare robots.txt, Content Signals, llms.txt, and RSL without treating access, use preferences, guidance, and terms as interchangeable.

Read the signal guide →

CrawlLedger 0.3.0 · first observed Aug 4, 2026 · latest observation