robots.txt & llms.txt across the Tranco top 1M

Domain explorer

Look up a single domain's robots.txt & llms.txt metrics, effective AI policy, inferred tech stack (labelled hypothesis), and crawl history. Read-only; one indexed lookup per query.

state.mn.us

Tranco rank
41,198
popularity rank
TLD
mn.us
state.mn.us
robots.txt records
1
across crawls
llms.txt records
1
across crawls

Live files on state.mn.us

Open the file as it exists on the host right now — may differ from what we measured. We link out rather than mirror (we store derived metrics, not raw bodies).

Captured samples

Snapshots we captured for the showcase set (may differ from the live file above). Large files are truncated; the true size is noted.

/robots.txt · 524 bytes
# Disallow everything until we want to expose the site to external search
# engines.
# 0000-1200 GMT is 6PM to 6AM here
# 4/14/14 Updated for DataExplorer to crawl all state sites.

User-agent: *
Disallow: /cgi-bin
Disallow: /cgi-sys
Disallow: /cd_upload/Search
Disallow: /law-library-stat/archive/
Disallow: /law-library-stat/briefs/
Visit-time: 0000-1200
Request-rate: 10

User-agent: Ultraseek
Disallow: /cgi-bin
Disallow: /cgi-sys
Disallow: /cd_upload/Search

User-agent: VSE/1.0
Disallow: /cgi-bin
Disallow: /cgi-sys


Effective AI / crawl policy (Google posture, §3.3)

CrawlPostureConsec. 5xxReasonAs of
CC-MAIN-2026-25 open 0 2xx: robots.txt present, no full disallow in effect

robots.txt

CrawlStatusBytesGroupsAllowDisallow SitemapsWildcard *Disallow-allAI tokensAI verdicts
CC-MAIN-2026-25 200 524 3 0 10 0 yes no

llms.txt

Collected by our own polite crawler, not Common Crawl — Common Crawl doesn't capture /llms.txt. See Methodology.

CrawlStatusVariantConformantSections LinksGenerator
CC-MAIN-2026-25 404

llms.txt data reflects adoption + conformance only, never consumption — publishing a file does not mean any AI reads it.

Inferred tech stack hypothesis

No tech inference for this domain.