CrawlRadar
OpenAI logo

OpenAI · training

Is GPTBot blocked on your site?

GPTBot is OpenAI’s training token for OpenAI model training, and whether it reaches you is decided by one line in your robots.txt. Blocking it keeps your pages out of future OpenAI model training. It does not remove you from ChatGPT Search, which crawls under OAI-SearchBot.

  • Not critical if blocked
  • GPTBot robots.txt token

By · Published · Last updated

What is GPTBot?

GPTBot is the crawler OpenAI uses to collect public pages that may be used to train its generative AI foundation models. OpenAI documents it separately from OAI-SearchBot, which decides whether a page can appear in ChatGPT Search, and from ChatGPT-User, which fetches a page when someone asks ChatGPT to. OpenAI states that each of these settings is independent of the others.

OperatorOpenAI
Purposetraining · OpenAI model training
robots.txt tokenGPTBot
In the user-agentcompatible; GPTBot/1.x; +https://openai.com/gptbot
CrawlRadar weight1 of the crawler-access weights, never flagged critical
Vendor docsOpenAI: Overview of OpenAI crawlers

How to allow GPTBot in robots.txt

Give GPTBot its own group with Allow: /. A named group overrides the User-agent: * group, so this works even when the wildcard disallows everything:

User-agent: GPTBot
Allow: /

How to block GPTBot in robots.txt

Replace Allow with Disallow. To block only part of the site, disallow that path instead of /:

User-agent: GPTBot
Disallow: /
  • Tokens are matched case-insensitively, but writing GPTBot exactly as the vendor spells it avoids surprises with stricter parsers.
  • Check your CDN and security plugins too: a firewall rule can block GPTBot without any line in robots.txt.

Should you block GPTBot?

Block GPTBot if you do not want your pages used to train OpenAI’s models: as long as OAI-SearchBot stays allowed, ChatGPT Search can still find and cite you. Allow it if you would rather future models learn your product and brand names without needing a live search. That effect is real but slow, arriving with new model releases rather than within weeks.

How to verify a request is really GPTBot

OpenAI publishes the IP addresses GPTBot uses at https://openai.com/gptbot.json. A request that claims to be GPTBot from any other address is someone borrowing the name, and should be treated like any other scraper. CrawlRadar’s collector checks every hit against that list and labels it verified or spoofed.

OpenAI’s other AI crawlers

Each OpenAI token is a separate robots.txt decision. Blocking one has no effect on the others.

TokenPurposeWhat blocking it costs
OAI-SearchBotsearch indexThis is the crawler that decides whether ChatGPT Search can cite you. Blocking it removes you from that citation pool entirely.
ChatGPT-Userlive answer fetchIt fetches your page while someone is waiting on an answer. Blocking it means ChatGPT may not open your page even when a user asks it to.

Is GPTBot visiting your site?

Your analytics will not tell you. AI crawlers do not run JavaScript, so a tag-based tool such as Google Analytics never records them. CrawlRadar’s collector reads your own request path and records each GPTBot visit, the page it fetched, and whether it came from OpenAI. See the full AI crawler directory for the other tokens CrawlRadar checks.

Frequently asked questions

Does blocking GPTBot remove my site from ChatGPT?

No. ChatGPT Search crawls under OAI-SearchBot, and live page fetches arrive as ChatGPT-User. Blocking GPTBot only opts your pages out of model training, which is why CrawlRadar reports a GPTBot block as an opt-out rather than a blocking issue.

How do I know a GPTBot request is really from OpenAI?

OpenAI publishes the IP addresses GPTBot crawls from at openai.com/gptbot.json. A request that says GPTBot but comes from another address is a scraper borrowing the name. CrawlRadar’s collector checks every hit against that list and labels it verified or spoofed.

See whether GPTBot actually visits you

robots.txt says who may crawl; only your own request log says who did. CrawlRadar’s collector records every AI crawler visit, verifies it against the vendor’s published ranges, and shows the pages each one fetched.

14-day free trial. No credit card.

Or start with the free AI crawler checker.