Reviewed 12 September 2026
Scan methodology and limitations
CanAIRead.me runs a technical homepage diagnostic. It observes access, rendering, content structure, structured data and basic search signals. It does not observe an AI model’s private index or guarantee that any system will mention or cite a page.
What the scan requests
The scanner resolves the submitted public hostname, rejects private or unsafe network targets, and tries HTTPS before a single HTTP fallback. It requests the homepage, /robots.txt, a declared or conventional XML sitemap, and /llms.txt. It also loads the homepage in headless Chromium to compare rendered text with text available in the initial HTML.
The scan is homepage-focused. It does not crawl every product, article, collection, CMS template or private page. Redirected and subresource behavior may differ from the initial homepage response.
Crawler directives
robots.txt is parsed for named agents including GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, Googlebot, Google-Extended, PerplexityBot, Bingbot, Applebot, Bytespider and CCBot. The report evaluates access to the homepage path and records the selected user-agent group and matching directive where one applies. A permissive robots.txt file does not prove that a CDN, WAF or origin server will accept the crawler’s real request.
HTML, rendering and schema
The raw response is inspected for title, description, H1 and heading structure, canonical URL, images without an alt attribute, robots meta directives, favicon and JSON-LD. JSON-LD types and selected identity properties are extracted, including nodes inside an @graph. Invalid JSON blocks are counted. A browser render comparison estimates how much visible text depends on JavaScript; it is not a Core Web Vitals test.
How the score is formed
Five category scores are combined: access 25%, readability 20%, understanding 25%, citation readiness 15%, and search foundations 15%. These weights are CanAIRead.me’s diagnostic model, not weights published by Google, OpenAI, Anthropic, Perplexity or another platform. “Citation readiness” means observable technical and content signals—not measured citations.
What a passing result does not mean
- • It does not guarantee crawling, indexing, ranking, inclusion in model training, retrieval, an AI mention or a citation.
- • It does not validate every page or every robots path rule.
- • It does not prove that schema is eligible for a Google rich result.
- • It does not measure brand demand, backlinks, factual authority or answer relevance.
- • A missing
llms.txtis not treated as a search requirement.
CMS detection
CMS identification uses signatures found in the returned homepage HTML, such as asset hosts, generator markers and platform attributes. A match is reported with confidence and evidence. If no reliable signature is present, the report says the platform is unconfirmed and keeps technical findings independent from platform-specific advice.
Privacy and retention
Scans are performed server-side against public URLs. Reports may be cached for up to 24 hours to avoid repeatedly requesting the same site, and recently checked hostnames may be shown publicly. Email delivery and analytics use the providers described in the privacy notice. Do not submit private, authenticated or confidential URLs.
Review cycle and corrections
The methodology is reviewed when scan logic or supported crawler guidance changes, and its date is updated only after a substantive review. Crawler behavior can change between reviews; the crawler directory links to primary documentation where available.