01 / Observe narrowly
One queued domain at a time, limited to public policy files in a reviewed 20-domain pilot.
CrawlLedger preserves what domains publish in robots.txt, Content Signals, llms.txt, and RSL licensing files—plus when it was observed, how it parsed, and how each record links to the one before it.
A normal checker tells you what a file says now. CrawlLedger keeps the observation metadata needed to distinguish a current result, an unchanged response, a parse failure, and a fetch skipped by robots policy.
One queued domain at a time, limited to public policy files in a reviewed 20-domain pilot.
Content hashes, bounded response metadata, parser versions, and per-domain signal chains remain attached to every event.
Pages use factual states such as observed, not observed, unverified, parse failed, or skipped—not legal conclusions.
Training crawlers, search crawlers, user-triggered fetchers, and control tokens are not interchangeable. CrawlLedger keeps their first-party descriptions and observed policy states separate.
Compare rules for GPTBot, ClaudeBot, and Applebot-Extended.
Inspect separate search controls for OAI-SearchBot, Claude-SearchBot, and PerplexityBot.
Understand the emerging Content-Signal fields for search, AI input, training, and reuse.
Only post-baseline byte changes that produced a parsed result are shown here. Changing error pages and routine 304 responses are excluded.
robots.txt · Observed and parsed
robots.txt · Observed and parsed
Each domain page is a canonical, crawlable landing page for the public metadata already available from the JSON API.
2 signal types · last observed Aug 15, 2026
2 signal types · last observed Aug 15, 2026
2 signal types · last observed Aug 15, 2026
3 signal types · last observed Aug 15, 2026
3 signal types · last observed Aug 15, 2026
The archive is useful whether you need to audit a crawler rule, cite a historical observation, monitor emerging licensing signals, or consume structured metadata from an API.
Paste a robots.txt file and evaluate the tracked AI crawler identifiers without triggering a network request or storing the input.
Compare robots.txt, Content Signals, llms.txt, and RSL without treating access, use preferences, guidance, and terms as interchangeable.
Read current observations or bounded history while raw source artifacts remain private.
CrawlLedger 0.3.0 · first observed Aug 4, 2026 · latest observation