Cookie text goes here

Now that we have your attention, we'd like you to know that we use cookies to enhance your browsing experience and analyse our traffic. By clicking 'Accept', you consent to our use of cookies.

Is Your Website Open to AI Crawlers?

Find Out in Seconds. 100% Free

Check which bots can access your site and uncover performance or search visibility issues that affect how your content is found online.

FAQ

This scanner emulates the user agents of popular generative AI tools and search engine crawlers to assess your website’s accessibility. It identifies which bots can successfully reach and index your content, and evaluates key technical SEO elements to help maximize your site's visibility in search results.

The scan audits a set of 32 sections and calculates a compatibility score (0-100%) based on the result:

 

The tool uses a Python3 backend application to crawl, collect and parse information displayed on the webpage. The following API is used:


The following libraries are used:

 

A user agent is a string of text that a browser, app, or bot sends to a web server to identify itself. It typically includes information about the software type, version, operating system, and sometimes the device. For example, a user agent might say: "Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/115.0.0.0 Safari/537.36"—which tells the server that the request is coming from a Chrome browser on Windows.

Web servers use user agents to tailor responses, such as serving mobile-friendly pages to smartphones or blocking known bots. In the case of web crawlers, the user agent often includes the name of the crawler (e.g., Googlebot or Bingbot) so websites can recognize and manage how their content is accessed and indexed.

A web crawler, also known as a spider or bot, is a software program designed to systematically browse and index the content of websites across the internet. It starts with a list of URLs and visits each one, extracting links and data, then follows those links to discover new pages. This process helps search engines like Google or Bing build massive indexes of web content, enabling users to find relevant information quickly through search queries.

Web crawlers are essential for keeping search engine databases up to date. They typically obey rules set in a site's robots.txt file, which tells them which pages they're allowed or disallowed from accessing. Some crawlers are specialized—for example, those used for price comparison, academic research, or monitoring website changes.

A Generative AI crawler operates similarly to traditional web crawlers but with a more advanced purpose: gathering data to train or enhance AI models. Instead of just indexing pages for search, these crawlers focus on collecting high-quality, diverse, and relevant content—such as text, images, or code—that can be used to improve the capabilities of generative AI systems.

These crawlers often use more sophisticated filtering and classification techniques to identify useful content while avoiding low-quality or irrelevant data. They may also respect copyright and content usage policies more strictly, especially when gathering data for commercial AI models. In some cases, they're designed to detect and avoid sensitive or restricted information, ensuring ethical and legal compliance in data collection.

A robots.txt file is a simple text file placed in the root directory of a website that tells web crawlers (also known as bots or spiders) which parts of the site they're allowed or disallowed to access. It's part of the Robots Exclusion Protocol, a standard that helps site owners manage how their content is indexed or scraped.

While robots.txt is a public file and not a security measure (bots can ignore it), it's a vital tool for SEO, content management, and controlling how your site interacts with search engines and AI crawlers.

Hypertext Transfer Protocol (HTTP) response status codes are issued by a server in response to a client's request made to the server.

All HTTP response status codes are separated into five classes or categories. The first digit of the status code defines the class of response, while the last two digits do not have any classifying or categorization role. There are five classes defined by the standard:

Made byThomas Granelund with ❤️ in Helsinki, Finland

Edit cookie consent