In my research, I've found five technical blockers that prevent a website from being seen by ChatGPT, Google, and other AI systems. These blockers stop the web crawlers that scan and index your website from accessing or reading its content. This post explains each of these blockers and how to test for them.
TL;DR
Check for these five potential blockers that could prevent web crawlers from reading and indexing your site:
Robots.txt must be set to allow web crawlers to crawl your site.
Your Content Delivery Network (CDN) cannot be blocking web crawlers.
JSON-LD should label the content on your site.
JavaScript rendering or interactive elements shouldn't hide content crawlers need to read.
llms.txt is recommended, since some LLMs are using it.
How to Check Robots.txt
Robots.txt is a file on your website that contains rules to allow or block specific web crawlers (robots or bots) from accessing your website (read more from Google).
You can see which web crawlers have access to your site by opening your robots.txt file.
Open a browser and go to the URL
https://yourwebsite/robots.txtLook for:
User-Agent: *
Allow: /Or look for specific website crawlers set as allowed or disallowed.
How to Check CDN (Content Delivery Network)
Your CDN might be set to block specific crawlers. A CDN has servers in different geographic areas that store website files so they can load faster. Popular CDNs are Cloudflare and Amazon Web Services. Some CDNs block AI bots by default.
You can check whether your CDN is blocking bots by using tools like aicrawlercheck.com, crawlercheck.com, or by checking your CDN's logs and settings.
How to Check for JSON-LD
JSON-LD (JavaScript Object Notation for Linked Data) is code added to the HTML on webpages on your site. It labels your content and tells web crawlers what's on the page.
You can check whether your site has JSON-LD by going to:
How to Check for JavaScript Rendering or Interactive Elements
JavaScript is code on a website that makes pages interactive. But some web crawlers don't run that code, which means content that depends on JavaScript running is invisible to them. For example, if you have important FAQ information in dropdowns or accordions on a webpage, that information might be invisible to the bots.
You can use these tools to check for problems preventing web crawlers from seeing content on your site:
How to Check for LLMS.txt
LLMS.txt is a proposed standard text file for your website that tells LLMs and AI agents what your website is about. Adoption by the LLMs is not consistent, but there's no penalty for adding it.
You can check whether your site has an LLMS.txt file by opening a browser window and entering your site with LLMS.txt at the end, like the example below:
Open
https://yourwebsite/LLMS.txt
Which web crawlers should have access to my site and should be listed in my robots.txt file?
You can allow all web crawlers to access your site by setting the User-Agent to User-Agent: *.
Here's a list of specific bots from OpenAI (ChatGPT), Google, Anthropic (Claude), and others:
Crawler | What it does |
|---|---|
OAI-SearchBot | Surfaces websites in ChatGPT's search features. |
GPTBot | Crawls content that may be used to train OpenAI's generative AI foundation models. |
ChatGPT-User | Visits a web page when a user asks ChatGPT a question that requires browsing it. |
ClaudeBot | Anthropic's crawler, collecting data to support Claude's training and search features. |
Meta-ExternalAgent | Meta's crawler, indexing content for AI training and direct platform features. |
Amazonbot | Amazon's bot, gathering web data to help power Alexa and related services. |
Bytespider | ByteDance's crawler, used for data collection and AI training. |
Googlebot | Google's primary crawler for indexing desktop and mobile web pages. |
Bingbot | Microsoft's crawler for the Bing search engine. |
Applebot | Apple's crawler, supporting Siri, Safari, and Spotlight. |
DuckDuckBot | The privacy-focused crawler for DuckDuckGo. |
If GPTBot is blocked but OAI-SearchBot isn't, does that stop training but still let ChatGPT cite the site in search?
Yes. GPTBot is training, OAI-SearchBot is search, and ChatGPT-User is the live browse when a user's question sends ChatGPT to a specific page. Blocking one doesn't block the others. Check robots.txt for all three separately.
The site loads fine in a browser. Why would a CDN block a bot at all?
Because bot protection is often set to be defensive by default. Not all bots are useful, and some can be harmful. You want to let the good bots in and keep the ones you don't want out.
More checklists like this go out as I write them.
