How to Check AI Crawler Access Beyond Your Homepage
Build a page-by-page AI crawler access check that separates robots.txt rules from HTTP responses and avoids mistaking access for visibility.

A homepage that loads for an AI crawler checker does not clear the rest of your site. Test a small set of important URLs individually: the homepage, a section page and deep pages that matter to readers. Record what robots.txt appears to permit separately from what the server returns. If a checker cannot make the request under the crawler identity you need to investigate, mark that test not run rather than calling the page accessible.
Choose pages and crawler identities
Start with URLs that might encounter different rules. For a publisher, that could mean the homepage, an articles section, a recent article and a useful older article under another path. Include a URL where you suspect a restriction. CrawlerCheck says it accepts an individual page URL, including deep content, so it is one option for assembling page-level findings rather than relying on a domain-level result (CrawlerCheck).
Make the checklist a grid: one row per URL and one column per crawler identity you intend to assess. Treat OAI-SearchBot and GPTBot as separate checks, not interchangeable names for an AI bot. Scrawl describes OAI-SearchBot as a search/retrieval crawler and GPTBot as a training crawler. Those are Scrawl’s characterizations; confirm current provider roles and identifiers in official documentation before making a provider-specific policy decision. If your goal is to investigate search access, a GPTBot-only result does not answer the OAI-SearchBot question.
Tool scope matters here. Scrawl says its robots.txt report names OAI-SearchBot, ChatGPT-User and GPTBot. CrawlerView lists GPTBot among its tested bots but does not list OAI-SearchBot. Siftly asks for a domain without a path, so its description does not establish that it tests each deep URL in your grid. Select a checker for the specific URL and identity you need; do not fill gaps in the grid by inference.
Read the policy before interpreting a request
For each URL and identity, first inspect the site’s robots.txt policy. Scrawl says it fetches robots.txt and labels named crawlers allowed, disallowed or not mentioned (Scrawl). “Not mentioned” is a report label, not a standalone permission verdict: review the applicable rules and the URL path before recording your policy finding. Keep the policy text or report alongside your interpretation so someone else can check it.
Next, test the URL’s HTTP response where a suitable page-level request check is available. A robots.txt report describes directives; it cannot show that a request succeeded or that a crawler obeyed them. Scrawl explicitly says its robots.txt check cannot guarantee crawler compliance. Conversely, an HTTP response alone does not tell you what the robots.txt policy says. Record these as two findings, even when they appear to agree.
A compact working record might look like this:
| Field | What to record |
|---|---|
| Target | Exact URL tested, including its path |
| Identity | Named crawler identity requested, or “not available” |
| Policy | Applicable robots.txt finding and checker |
| Response | Observed status, destination URL if redirected, and whether the intended page was returned |
| Conditions | Checker, request identity, network context and test time |
The last row prevents a common misread. Glippy’s live check says it requests the homepage under multiple identities from one address, changing the User-Agent between requests. That comparison can reveal different responses under those conditions, but it cannot clear a deep article path. Glippy also says its live check runs from an edge-network address that some firewalls refuse; a refusal there may warrant a repeat check from another network, not an immediate conclusion that the provider’s crawler is blocked. Its live check does not inspect robots.txt, meta tags or client-side rendering, so it cannot replace the policy review.
Investigate a mismatch, then repeat the affected test
Suppose your checklist shows that a deep article is not disallowed under your reading of robots.txt, but a simulated bot request receives a 403 while a normal browser request receives a 200. Investigate the firewall or server rules for that URL and request identity before changing robots.txt. CrawlerCheck discusses this pattern for a simulated Googlebot request and identifies a WAF or server configuration as a possible cause (CrawlerCheck); its example is not proof of the cause on your site or a result for OAI-SearchBot.
Preserve intentional restrictions. For an unintended restriction, make a targeted change and repeat the affected URL test under the recorded conditions. Call the access problem resolved only when that URL returns the intended page in the repeated check. If it redirects, inspect the destination rather than treating the initial URL’s response as the whole finding. CrawlerView says it follows specified redirects to the final destination (CrawlerView).
Finally, keep the scope of the finding honest. A request sent with a crawler-like identity is a simulated request, not a verified visit by that provider. A successful fetch establishes a response under the tested conditions; it does not establish indexing, retrieval, a citation or a referral. Meta directives, rendered content and other eligibility questions deserve their own checks after server access. For a broader account of what different checks can establish, see What Can AEO and GEO Tools Actually Measure?.




