WavroWavro Labs · Free tool

AI Bot Accessibility Checker

Is your site letting the AI crawlers in?

Check any URL against 14 AI crawler user-agents including GPTBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot. Cross-checks your robots.txt rules against real crawler access, catching CDN and WAF blocks that robots.txt-only parsers miss. Free, no login, no signup.

Used it already? Read reviews

We check your public robots.txt and verify live access. Nothing is stored or logged.

14
User-agents checked
Across 10 AI platforms
60–90 s
Average check time
3 signals
Chat · Train · Crawl
Free
No signup required
What this tool checks

14 AI crawlers, in one scan

We parse the User-agent and Disallow directives in your robots.txt for every major AI crawler on the public web, then cross-check those rules against real crawler access to see how your origin (or CDN/WAF) actually responds. For each platform you get three signals:

  • ChatQuery-time agent (fetches pages for AI answers), from robots.txt.
  • TrainTraining-time crawler, from robots.txt.
  • CrawlVerified access: does the crawler actually reach the page? Catches CDN and WAF blocks that robots.txt can't.
Why this matters

Blocking AI crawlers is a business decision, not a default

You disappear from AI answers

If GPTBot / ClaudeBot / PerplexityBot cannot fetch your pages at query time, your brand cannot be cited when users ask AI about your topic, even if you rank #1 on Google.

Training-time blocks are separate

Blocking a training-time crawler doesn't stop that AI from referencing your site in real time. Many sites confuse the two and lose citations by accident.

CDN defaults may be biting you

Cloudflare and other CDNs added "block AI crawlers" toggles in 2024. Many sites enabled them without realizing they were opting out of AI Overviews and Perplexity citations.

The crawler catalog

Which AI bots we check for

GPTBot
OpenAI

Trains ChatGPT / GPT models.

Official docs
OAI-SearchBot
OpenAI

Powers ChatGPT Search real-time citations.

Official docs
ChatGPT-User
OpenAI

Fetches on-demand when a user asks a question in ChatGPT.

Official docs
ClaudeBot
Anthropic

Trains Claude on public web content.

Official docs
Claude-Web
Anthropic

Fetches at query time for Claude answers.

Official docs
Claude-User
Anthropic

User-triggered fetches when someone asks Claude about a page.

Official docs
anthropic-ai
Anthropic

Legacy Anthropic user-agent, still honored in robots.txt.

Official docs
PerplexityBot
Perplexity

Crawls for Perplexity search citations.

Official docs
Perplexity-User
Perplexity

User-triggered fetches at answer time.

Official docs
Google-Extended
Google

Opt-out token for Gemini + Vertex AI training. Does not affect Search.

Official docs
Googlebot
Google

Fetches pages that feed AI Overviews, AI Mode, and Gemini answers.

Official docs
bingbot
Microsoft

Powers Copilot retrieval. Microsoft publishes no dedicated AI training bot.

Official docs
Amazonbot
Amazon

Crawls for Rufus, Amazon’s shopping assistant.

Official docs
CCBot
Common Crawl

Feeds nearly every LLM (GPT, Claude, LLaMA) indirectly via the Common Crawl dataset.

Official docs
FAQ

AI bot access: common questions

Should I block AI crawlers?

For most content sites the answer is no. Blocking real-time crawlers like PerplexityBot, ChatGPT-User, or Claude-User means your brand cannot be cited when users ask AI about your topic. Blocking training-time crawlers (GPTBot, ClaudeBot, CCBot) is a separate philosophical question about copyright and revenue share.

What is the difference between chat and train crawlers?

Chat crawlers (also called query-time or user-triggered) fetch your page in real time when someone asks the AI a question. Train crawlers scrape your site to build the model itself. Blocking one does not block the other, and blocking training does not stop AI from citing you at answer time.

How do I allow AI crawlers?

Explicitly allow them in your robots.txt with User-agent: GPTBot / Allow: /. Same pattern for each user-agent listed above. If you use Cloudflare or another CDN, also check its dashboard for any "block AI bots" toggle that overrides robots.txt.

Does robots.txt actually stop AI crawlers?

For compliant operators (OpenAI, Anthropic, Perplexity, Google), yes, they respect it. For non-compliant scrapers, no. But robots.txt is the industry-standard signal and the courts have consistently treated ignoring it as bad faith. If you want a hard block, IP-level firewalls or a WAF are the enforcement layer.

Do you store the URLs I check?

No. This tool checks your robots.txt and verifies crawler access at request time, then returns the result. Nothing is persisted. If you want a persistent record of your AI visibility over time, run a full Wavro audit instead: it saves the result and can be shared.

WavroBeyond bot access

Get a free AI-visibility audit

Bot access is one of 25+ checks Wavro scores your site on across AEO, GEO, and LLMO. See exactly where you rank, plus copy-paste JSON-LD, alt text, and schema fixes for every failing check.

  • 25+ checks across AEO / GEO / LLMO
  • Ready-to-paste JSON-LD + alt text
  • Free, no signup required