Can a crawler reach the homepage?
Weighted most. A firewall challenge, an error or a redirect loop stops everything else.
- Homepage answers with a normal page
- No bot challenge in front of it
- No noindex on the homepage
Free tool
Type a domain and see, in about 20 seconds, what search and AI crawlers can reach and read on its homepage. The results need no sign-up.
What the check covers
Free, no sign-up. Results in about 20 seconds; the fix for each problem unlocks with your email.
13
Crawlers checked in robots.txt
~20 sec
For a full check
$0
Results, no sign-up
Raw HTML
Read the way crawlers do
What it checks
The checker answers them from your server's real response, in the order a crawler meets them.
Weighted most. A firewall challenge, an error or a redirect loop stops everything else.
Search and AI search crawlers are scored. Training crawlers are listed but not scored.
Many AI crawlers do not run JavaScript. If the text arrives later, they may see an empty page.
Small signals that tell every engine what the page is and which address is the real one.
Structured data that names the business or person, and links the profiles that confirm it.
A sitemap points crawlers to every page you want found, not just the homepage.
The 13 crawlers
Blocking a training crawler and blocking a search crawler are very different decisions. Only search and AI search crawlers count towards your score.
Crawler
Company
What it does
Scored
Googlebot
Google Search, including AI Overviews and AI Mode
Search
Bingbot
Microsoft
Bing search, which Copilot draws on
Search
OAI-SearchBot
OpenAI
Shows sites in ChatGPT search results
AI search
ChatGPT-User
OpenAI
Fetches a page when a ChatGPT user asks for it; OpenAI says robots.txt may not apply
User fetch
GPTBot
OpenAI
Collects content for training OpenAI models
Training
PerplexityBot
Perplexity
Indexes sites for Perplexity answers
AI search
Perplexity-User
Perplexity
Fetches a page when a Perplexity user asks for it; Perplexity says it generally ignores robots.txt
User fetch
Claude-SearchBot
Anthropic
Indexes sites for Claude search results
AI search
Claude-User
Anthropic
Fetches a page when a Claude user asks for it
User fetch
ClaudeBot
Anthropic
Collects content for training Claude models
Training
Google-Extended
Controls use in Gemini training and grounding in Gemini Apps (not a crawler)
Training
Applebot-Extended
Apple
Controls use in Apple AI training (not a crawler)
Training
CCBot
Common Crawl
Open web archive many AI models are trained on
Training
How a check runs
Any public site
Yours or a competitor's. The checker fetches public pages only and never logs in or submits forms.
About 20 seconds
Homepage, robots.txt, sitemap and llms.txt are requested from a server, the way a crawler would see them.
Free
A score, every check with pass, warning or fail, and the crawler-by-crawler robots.txt result.
Optional
Add your email to see how to fix each problem. I get a copy and may follow up once, never more.
Where it stops
The checker answers one question well: can crawlers reach and read your homepage? These need a person and a full audit.
A $149 audit covers one service area across the whole site, written up with every problem ranked and the fix spelled out.
Checker questions
Want the difference between SEO, GEO and AEO first? Read the comparison.
It fetches your homepage, robots.txt, sitemap and llms.txt from a server, the way a crawler does. It reports which of 13 search and AI crawler tokens your robots.txt allows, how much text is in the HTML before JavaScript runs, your title, canonical and language tags, structured data and sameAs links, noindex directives and firewall challenges.
Each check has a weight. Reachability, indexability, search crawler access and content in the HTML count most; small tags count least. A warning earns half the points of a pass. If the homepage cannot be reached at all, the score is capped at 20, because nothing else can work until it can.
Because on current evidence neither affects whether Google Search, ChatGPT search or Perplexity can show your site. The checker still lists them so you can decide, and notes that Google-Extended also covers grounding in Gemini Apps.
No. It checks whether AI crawlers can reach and read your site, which has to come first. Testing what assistants actually say takes a fixed set of prompts run across several engines, which is part of the GEO audit.
If the checker gets a challenge page or a 403, your firewall or bot protection may be doing the same to AI crawlers. Check your Cloudflare, host or security plugin settings and allow the verified crawlers you want.
Yes, and you can run it as often as you like. Adding your email shows the fix for each problem. If you want the fixes done for you, that is what the $599 sprint is for.
After the check
Send the domain you checked and the line that worried you. I will explain what it means for your site, and whether it is worth fixing.
I reply by email, usually within one working day.