The short answerYour robots.txt file controls which AI crawlers can access your site. Many South African websites are accidentally blocking ChatGPT and Perplexity.
The short answer
Your robots.txt file controls which AI crawlers can access your site. Many South African websites are accidentally blocking ChatGPT and Perplexity.
robots.txt and AI Crawlers: Why It Matters
Direct answer: robots.txt is a plain-text file at the root of your website that tells web crawlers which pages they can and cannot access. It has been a standard part of the web since 1994 and is respected by all major crawlers: including the AI crawler bots that power ChatGPT, Perplexity, and other AI answer engines.
Current as of 31 May 2026: This article has been reviewed for the 2026 South African AI, SEO, and automation market. Pricing, platform capabilities, Google rich-result rules, and AI model features change quickly, so verify live vendor documentation before procurement. For privacy and data handling, use the Protection of Personal Information Act as the baseline; for search and structured-data implementation, use Google Search Central.
The problem: many South African websites have robots.txt configurations that accidentally block AI crawlers, either through overly broad bot-blocking rules or by explicitly blocking AI user agents. If an AI crawler cannot access your pages, those pages will never be cited in AI-generated answers: regardless of how good your content is.
This guide explains the AI crawler landscape, how to audit your robots.txt, and how to configure it correctly for maximum AI visibility.
The Major AI Crawlers and Their User-Agent Strings
Knowing the crawlers by name is essential for targeted robots.txt configuration:
| AI System | Crawler User-Agent | Operator |
|---|---|---|
| ChatGPT (browsing) | OAI-SearchBot | OpenAI |
| ChatGPT (training) | GPTBot | OpenAI |
| Claude (Anthropic) | ClaudeBot | Anthropic |
| Perplexity | PerplexityBot | Perplexity AI |
| Google AI Overviews | Googlebot | |
| Apple Intelligence | Applebot-Extended | Apple |
| Meta AI | Meta-ExternalAgent | Meta |
| Microsoft Copilot | Bingbot | Microsoft |
Note that some AI systems reuse existing crawler identities (Google AI Overviews uses standard Googlebot; Copilot uses Bingbot). Blocking these would also impact Google and Bing rankings. Which is almost never desirable.
How AI Crawlers Get Accidentally Blocked
Scenario 1: The "Disallow Everything" Rule
The most common accidental block is a catch-all disallow rule:
User-agent: * Disallow: /
This blocks ALL crawlers including every AI bot. This is sometimes used temporarily during site development and forgotten. It can silently kill your GEO visibility.
Scenario 2: Security Plugin Bot Blocking
WordPress security plugins (Wordfence, iThemes Security, Sucuri) often include bot-blocking features that add broad user-agent restrictions to robots.txt or .htaccess. These rules frequently catch legitimate AI crawlers alongside malicious bots.
Check: if your robots.txt includes long lists of blocked user agents generated by a plugin, review whether AI crawlers are included.
Scenario 3: Explicit GPTBot Blocking
In 2023, many websites added Disallow: / rules for GPTBot to prevent OpenAI from using their content for model training. This is a valid concern for publishers worried about AI training data, but blocking GPTBot also blocks ChatGPT's web browsing capability. These are different functions with different business implications.
Scenario 4: CDN or WAF Rules
Content Delivery Networks (CDNs) and Web Application Firewalls (WAFs) sometimes block crawlers at the infrastructure level, bypassing robots.txt entirely. These blocks appear in server logs as 403 errors for specific user agents. Your robots.txt may look fine while the CDN is silently blocking AI crawlers.
Auditing Your Current robots.txt
Step 1: View your file Navigate to yourdomain.co.za/robots.txt in a browser. The file should load as plain text.
Step 2: Look for these red flags:
- Any
Disallow: /rule underUser-agent: * - Explicit rules for
GPTBot,OAI-SearchBot,PerplexityBot, orClaudeBot - Very long lists of blocked user agents (sign of aggressive bot-blocking plugin)
- Missing file entirely. This is fine (no robots.txt means all crawlers allowed), but you should create one with proper directives
Step 3: Test with the Google Search Console robots.txt tester Search Console has a robots.txt testing tool that lets you enter specific URLs and user agents to see if they are allowed or blocked.
Step 4: Check server logs If you have server log access, search for 403 or 401 responses to these user agents: OAI-SearchBot, PerplexityBot, ClaudeBot. These indicate infrastructure-level blocking that robots.txt alone would not reveal.
Our GEO audit tool automatically checks your robots.txt for AI crawler blocks as part of its assessment.
The Recommended robots.txt Configuration for SA Businesses
Here is a well-structured robots.txt that protects legitimate business interests while maximizing AI visibility:
# Standard crawlers
User-agent: * Allow: / Disallow: /admin/ Disallow: /wp-admin/ Disallow: /private/ Disallow: /staging/ Sitemap: https://www.yourdomain.co.za/sitemap.xml
# OpenAI: allow browsing, restrict training if desired
User-agent: OAI-SearchBot Allow: /
# GPTBot (OpenAI training data): choose your stance:
# Option A: Allow (helps with ChatGPT knowledge)
User-agent: GPTBot Allow: /
# Option B: Block training but allow browsing
# User-agent: GPTBot
# Disallow: /
# Perplexity
User-agent: PerplexityBot Allow: /
# Anthropic (Claude)
User-agent: ClaudeBot Allow: /
# Apple Intelligence
User-agent: Applebot-Extended Allow: /
# Google (includes AI Overviews)
User-agent: Googlebot Allow: /
Key decisions in this configuration: 1. The default rule allows all crawlers except specified private paths 2. AI crawlers are explicitly allowed (overrides any catch-all blocks from plugins) 3. GPTBot has a documented choice between allowing and blocking: each business must decide
The GPTBot Decision: Training vs Browsing
OpenAI uses two separate crawlers:
- GPTBot: crawls content for model training data
- OAI-SearchBot: crawls content for ChatGPT's real-time web browsing
Blocking GPTBot prevents your content from being used to train future ChatGPT models but does NOT prevent ChatGPT from browsing your site for real-time answers. Blocking OAI-SearchBot prevents real-time citation in ChatGPT responses.
For most South African businesses, allowing both is the right choice. The risk of ChatGPT incorporating your content into its training is minimal for business websites, and the benefit of being browsable for real-time responses is significant.
Publishers, journalists, and original content creators have stronger reasons to consider blocking GPTBot specifically while keeping OAI-SearchBot open.
Protecting Sensitive Pages Without Blocking AI Visibility
Most robots.txt AI visibility issues come from over-blocking. The right approach is to block specific paths rather than all crawlers:
Block appropriately:
/admin/. Your website admin panel/wp-admin/: WordPress admin/private/: internal documents/staging/: development or staging content- Login and checkout pages (personal/transactional pages)
Never block:
- Service pages
- Blog and resource content
- About and team pages
- Contact pages
- Product or pricing pages
Checking If Your Changes Work
After updating robots.txt:
- Wait 24-48 hours: crawlers may have a cached copy of your old robots.txt 2. Use Google Search Console's robots.txt tester to verify specific paths
- Monitor server logs over the following weeks for AI crawler visits 4. Re-run the GEO audit to confirm the technical block is resolved
Note: Fixing robots.txt does not immediately produce AI citations. Crawlers need to re-index your content first (typically 2-4 weeks for active AI crawlers), then your content needs to meet quality and relevance thresholds for citation.
For a complete technical GEO assessment including robots.txt, llms.txt, schema markup, and content quality, use our free GEO audit tool or engage our GEO optimization service for expert implementation support.
Related Resources:
- Free GEO Audit Tool
- GEO Optimization Service
- llms.txt: The Complete Setup Guide for South African Websites
- GEO Audit Checklist for South African Businesses
Our experience at Smart AI Solutions shows that South African businesses see the strongest ROI when they start with a single, well-defined automation use case.
Further Reading:
Frequently Asked Questions
Should I block GPTBot from crawling my website?
This depends on your business type. For most South African businesses, allowing GPTBot is beneficial. It may improve ChatGPT's knowledge of your company. However, if you are a publisher or content creator concerned about your original work being used for AI training without compensation, blocking GPTBot while keeping OAI-SearchBot open is a reasonable compromise.
How do I know if my site is currently blocking AI crawlers?
Check yourdomain.co.za/robots.txt for Disallow rules under User-agent: * or specific AI bot names. Also run our free GEO audit at /tools/geo-audit which automatically checks your robots.txt for AI crawler blocks.
If I have no robots.txt file, does that block AI crawlers?
No. A missing robots.txt means all crawlers are allowed to access all public pages. This is the maximum-access default. You should create a robots.txt to explicitly manage access, but a missing file does not block AI crawlers.
Can I block AI crawlers for some pages but not others?
Yes. You can add path-specific Disallow rules for specific AI crawler user agents. For example, you could allow Perplexity to crawl everything except your /case-studies/ directory if those contain confidential client information.
Does blocking AI crawlers protect my content from being copied?
robots.txt is a courtesy protocol: well-behaved crawlers respect it, but it is not a technical enforcement mechanism. Legitimate AI companies like OpenAI, Anthropic, and Google respect robots.txt. It does not protect against bad actors who ignore the standard.
How often do AI crawlers check for robots.txt updates?
Major AI crawlers check robots.txt on every crawl session, typically every few days to few weeks for active websites. After updating your robots.txt, expect crawlers to pick up the changes within 1-2 weeks.
Does Google AI Overviews use a different crawler than regular Googlebot?
Google AI Overviews sources content using the standard Googlebot crawl index rather than a separate AI-specific crawler. This means your existing Google indexation is the primary input for AI Overview inclusion: maintaining good Googlebot accessibility is essential for both Google rankings and AI Overview appearances.




