GPTBot
OpenAI · Training crawler · Updated 2026-07-19
GPTBot is OpenAI’s training crawler — it crawls the web to collect content for model training. Blocking it opts your content out of future models but does not affect live answers.
User-agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbotSource: vendor-documented at developers.openai.com/api/docs/bots.
Does GPTBot respect robots.txt?
Controllable via robots.txt. OpenAI’s wording: “Disallowing GPTBot indicates a site’s content should not be used in training generative AI foundation models.”
Allow or block GPTBot
# Allow (recommended if you want GPTBot traffic)
User-agent: GPTBot
Allow: /
# Block
User-agent: GPTBot
Disallow: /Verifying it’s really GPTBot
OpenAI publishes GPTBot IP ranges at openai.com/gptbot.json. User-agent strings can be spoofed by anyone — identity claims are only trustworthy when the source IP matches the operator’s published ranges.
Fetch the published ranges and check a suspect IP against them:
curl -s https://openai.com/gptbot.jsonWhat site owners should know
2026 robots.txt surveys consistently rank GPTBot the most-blocked AI crawler (CCBot rivals or exceeds it among news publishers) — but note what blocking does and does not do: it opts you out of training; it does not remove you from ChatGPT search answers (that’s OAI-SearchBot) or stop live user fetches (ChatGPT-User).
Scan your site to see how your robots.txt, content negotiation, and discovery files treat AI agents — including whether you’re blocking agents you meant to allow. Background: the robots.txt guide, how AI agents browse, and agent readiness.
All AI agent user-agents
ClaudeBot · Claude-User · Claude-SearchBot · GPTBot · OAI-SearchBot · ChatGPT-User · PerplexityBot · Perplexity-User · Google-Extended · Google-Agent · Google-CloudVertexBot · Meta-ExternalAgent · Meta-ExternalFetcher · MistralAI-User · Applebot · Applebot-Extended · Amazonbot · bingbot · DuckAssistBot · CCBot · Bytespider