In my research, I've found five technical blockers that prevent a website from being seen by ChatGPT, Google, and other AI systems. These blockers stop the web crawlers that scan and index your website from accessing or reading its content. This post explains each of these blockers and how to test for them.

TL;DR

Check for these five potential blockers that could prevent web crawlers from reading and indexing your site:

  • Robots.txt must be set to allow web crawlers to crawl your site.

  • Your Content Delivery Network (CDN) cannot be blocking web crawlers.

  • JSON-LD should label the content on your site.

  • JavaScript rendering or interactive elements shouldn't hide content crawlers need to read.

  • llms.txt is recommended, since some LLMs are using it.

How to Check Robots.txt

Robots.txt is a file on your website that contains rules to allow or block specific web crawlers (robots or bots) from accessing your website (read more from Google).

You can see which web crawlers have access to your site by opening your robots.txt file.

  • Open a browser and go to the URL https://yourwebsite/robots.txt

  • Look for:

User-Agent: *
Allow: /
  • Or look for specific website crawlers set as allowed or disallowed.

How to Check CDN (Content Delivery Network)

Your CDN might be set to block specific crawlers. A CDN has servers in different geographic areas that store website files so they can load faster. Popular CDNs are Cloudflare and Amazon Web Services. Some CDNs block AI bots by default.

You can check whether your CDN is blocking bots by using tools like aicrawlercheck.com, crawlercheck.com, or by checking your CDN's logs and settings.

How to Check for JSON-LD

JSON-LD (JavaScript Object Notation for Linked Data) is code added to the HTML on webpages on your site. It labels your content and tells web crawlers what's on the page.

You can check whether your site has JSON-LD by going to:

How to Check for JavaScript Rendering or Interactive Elements

JavaScript is code on a website that makes pages interactive. But some web crawlers don't run that code, which means content that depends on JavaScript running is invisible to them. For example, if you have important FAQ information in dropdowns or accordions on a webpage, that information might be invisible to the bots.

You can use these tools to check for problems preventing web crawlers from seeing content on your site:

How to Check for LLMS.txt

LLMS.txt is a proposed standard text file for your website that tells LLMs and AI agents what your website is about. Adoption by the LLMs is not consistent, but there's no penalty for adding it.

You can check whether your site has an LLMS.txt file by opening a browser window and entering your site with LLMS.txt at the end, like the example below:

  • Open https://yourwebsite/LLMS.txt

Which web crawlers should have access to my site and should be listed in my robots.txt file?

You can allow all web crawlers to access your site by setting the User-Agent to User-Agent: *.

Here's a list of specific bots from OpenAI (ChatGPT), Google, Anthropic (Claude), and others:

Crawler

What it does

OAI-SearchBot

Surfaces websites in ChatGPT's search features.

GPTBot

Crawls content that may be used to train OpenAI's generative AI foundation models.

ChatGPT-User

Visits a web page when a user asks ChatGPT a question that requires browsing it.

ClaudeBot

Anthropic's crawler, collecting data to support Claude's training and search features.

Meta-ExternalAgent

Meta's crawler, indexing content for AI training and direct platform features.

Amazonbot

Amazon's bot, gathering web data to help power Alexa and related services.

Bytespider

ByteDance's crawler, used for data collection and AI training.

Googlebot

Google's primary crawler for indexing desktop and mobile web pages.

Bingbot

Microsoft's crawler for the Bing search engine.

Applebot

Apple's crawler, supporting Siri, Safari, and Spotlight.

DuckDuckBot

The privacy-focused crawler for DuckDuckGo.

If GPTBot is blocked but OAI-SearchBot isn't, does that stop training but still let ChatGPT cite the site in search?

Yes. GPTBot is training, OAI-SearchBot is search, and ChatGPT-User is the live browse when a user's question sends ChatGPT to a specific page. Blocking one doesn't block the others. Check robots.txt for all three separately.

The site loads fine in a browser. Why would a CDN block a bot at all?

Because bot protection is often set to be defensive by default. Not all bots are useful, and some can be harmful. You want to let the good bots in and keep the ones you don't want out.

More checklists like this go out as I write them.

Recommended for you

View all
caret-right