GenGrowth

GEO tools

Page Citability Check

Paste one URL to inspect crawler access, answer-structure patterns and raw/render evidence. Base checks need no account or language model. Optional AI semantic review requires sign-in and a separate explicit request.

One page · Base checks need no sign-in

Input

One page, not a whole site. The page has to return 200 and be HTML.

Fill this in and the opening-answer row is checked against it. Leave it empty and that row reports itself as not applicable rather than failing.

Base checks do not call a language model; optional AI review is requested separately.

What this check does not do

  • Two fresh Chromium contexts compare the same HTML with JavaScript off and on, within 12 seconds, 40 GET requests and 2 MB. Public images/media/fonts are intentionally omitted; private targets, workers, sockets, downloads and navigation are blocked. Original HTTP CSP is not preserved; a stricter safety policy is applied. This measures controlled body text, not full image/layout fidelity or every browser. Incomplete captures have no ratio.
  • It reads one page plus that host's robots.txt and llms.txt. It does not crawl the site and does not fetch the canonical target.
  • It says nothing about ranking, traffic or whether any model has actually cited the page. It describes the page, not the market.
  • ClaudeBot and GPTBot training access, Google-Extended Gemini training/grounding controls and /llms.txt are outside the generic pass count. Claude-SearchBot search and Claude-User user-directed fetching are not measured. Google Search crawling, AI Overviews and AI Mode use separate Search controls.
  • Three rows use pattern matching, not semantic or factual verification: opening-answer markers, quantified conditions, and numeric source cues. Review their evidence before acting.
  • It reads a bounded slice of the page. When a document exceeds it, every check that would report an absence says it could not read instead.

Questions about this check

Does passing every row mean a model will cite the page?
No. These rows describe access and shape: whether a crawler is allowed in, whether the copy is in the HTML, and whether a claim can be lifted out. Whether any model actually answers with your page depends on the question asked and on what else exists on the web. This tool never claims otherwise.
Why are ClaudeBot and GPTBot shown but not counted?
ClaudeBot and GPTBot collect training material. Blocking them does not establish a retrieval failure. Anthropic uses Claude-SearchBot for search and Claude-User for user-directed fetching; neither is measured here. Google-Extended additionally controls Gemini grounding in Gemini Apps and Vertex AI, not Google Search inclusion or ranking, including AI Overviews and AI Mode. These controls remain separate from a generic citation score.
Why does a missing robots.txt pass?
RFC 9309 treats an unavailable robots.txt as full allowance, so a site with no file is not blocking anyone. An unreachable robots.txt, such as a 500, is different: the site's rules are unknown, and that row says so rather than guessing.
The page looks fine in my browser but the HTML row failed.
The check compares the exact fetched HTML with an isolated Chromium render. If most copy appears only after scripts execute, a non-rendering reader receives less content. The report shows both captures; blocked resources, truncation or a missing renderer make the comparison unknown instead of passing it.

Where to go next

This check answers one question about one page. The rest is elsewhere.