home / knowledge base / ai visibility checklist
knowledge base

AI Visibility Checklist: Seven Checks You Can Run Yourself Today_

Whether an AI assistant mentions your business only partly depends on how you write. It mainly depends on seven technical things you can check yourself in an afternoon. Below, each check explains what to do and what a good outcome looks like.

the_core

Why you need to check this yourself right now

On August 21, 2026, Cloudflare announced Bot Preference Sync: a feature that automatically syncs your robots.txt with the AI bot policy set in the Cloudflare dashboard. Your own rules stay in place, but something gets layered on top. The feature rolls out to all plans starting the week of August 24, and is on by default for new customers.

Well intentioned. But it means the file crawlers fetch from your domain is no longer necessarily the file you wrote.

At the same time, this is the first year you can actually measure the outcome. Microsoft took the AI Performance report in Bing Webmaster Tools out of preview on February 11, 2026: it shows how often your pages get cited in Copilot, and for which questions.

check_01

Read your robots.txt the way a crawler receives it

Compare the file in your browser to the file you manage, and check which bots are named in it.

good outcome: both match, or you know which layer is causing the difference
check_02

Know which bot you're blocking, and why

Training, search and citation are separate crawlers with separate consequences; blocking GPTBot doesn't remove you from the answers.

good outcome: a deliberate choice; opting out of training is fine, out of search and citation never
check_03

Are you in Bing, and are you being cited there

Count your pages with site:yourdomain.com, compare that to your sitemap and open the AI Performance report in Bing Webmaster Tools.

good outcome: every sitemap page shows up in Bing and the AI report isn't empty
check_04

Give the signal yourself when you publish

Check that the IndexNow key file is in your root folder and that your publishing process really sends the ping.

good outcome: you can point to the ping and the last call returned a 200 or 202
check_05

Is your text actually in the source code

View the page source and use ctrl-F to find a literal sentence, on your service pages too.

good outcome: your text is in the source code of every page you want to be found on
check_06

Schema.org: hygiene, not leverage

Check that every important page has a JSON-LD block that matches the real content.

good outcome: a valid schema block on every important page, without expecting more from it
check_07

llms.txt: put it up, expect nothing from it

A summary and signpost for language models in your root folder; no major provider has committed to using it.

good outcome: it's up, it took half an hour, and you expect nothing from it
the seven checks at a glance: what you check and when it's good, worked out per check below
check_01

Read your robots.txt the way a crawler receives it

Open yourdomain.com/robots.txt in a private window and compare it to the file as it exists in your project or CMS. If the browser shows more than that, something else is writing to it too: your CDN, a security plugin or your hosting. Then check who is named in it. The names that matter today are GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended and Applebot-Extended.

Good outcome: the file in your browser matches the file you manage, or you know exactly which layer is causing the difference and why.

check_02

Know which bot you're blocking, and why

This is where things often go wrong, because "blocking AI" isn't a single switch but a series of separate choices. OpenAI uses three crawlers, each with its own purpose: GPTBot fetches material for training, OAI-SearchBot builds the index that ChatGPT's citations are drawn from, and ChatGPT-User fetches a page when a conversation calls for it.

Block GPTBot and you opt out of training, but you don't disappear from the answers. Block OAI-SearchBot and you disappear from the citations without stopping training at all. Google has the same dynamic in a different form: turning off Google-Extended doesn't remove you from AI Overviews, because those are tied to the regular search index. Cloudflare uses exactly that distinction: Search, Agent and Training.

Good outcome: you made a deliberate choice. For most SMEs that means: opting out of training is fine if you want to, but never out of search and citation.

check_03

Are you in Bing, and are you being cited there

Search site:yourdomain.com in Bing, count the pages and compare that to your sitemap. Then submit your site to Bing Webmaster Tools and open the AI Performance report. If everything's at zero while your pages are neatly indexed, that's a content problem. If your pages aren't indexed at all, it's the reverse.

Good outcome: every page from your sitemap shows up in Bing, and the AI report isn't empty.

check_04

Give the signal yourself when you publish

IndexNow is the protocol that lets you tell a search engine directly that something new is up, instead of waiting for a crawler to stop by. It requires a key file in your site's root folder and a call on every publish. Because Bing is the engine behind Copilot, it carries through into AI answers.

Check that the key file is actually there, and that your publishing process really sends the ping. A script that fails silently is the classic trap here: nothing visibly changes, so nobody notices.

Good outcome: you can point to exactly where the ping sits in your publishing process, and the last call returned a 200 or 202.

check_05

Is your text actually in the source code

Many AI crawlers run little or no JavaScript. If your text only gets loaded in by a script, a crawler like that sees an empty page. Google handles this more gracefully, which is exactly why you can rank fine in Google and be nowhere in AI answers.

The test takes ten seconds: right-click, choose "view page source" and use ctrl-F to search for a literal sentence from your own text. Do this on your service pages too, not just the homepage. Heavy page builders are the usual culprit here.

Good outcome: you find your text in the source code, on every page you want to be found on.

check_06

Schema.org: hygiene, not leverage

Structured data through schema.org tells a machine what's on your page: Organization, Article or BlogPosting, FAQPage, BreadcrumbList. Check that every important page has a JSON-LD block like this, and that it actually matches the real content.

Be honest about what it actually delivers. Cited pages carry JSON-LD roughly three times as often, and that number gets used as proof all the time. But a study of 1,885 pages that added JSON-LD, compared against roughly 4,000 pages that didn't make that change, found no clear increase in citations. It correlates, it doesn't cause.

Good outcome: every important page has a valid schema block that matches its content. Expecting more than that isn't realistic.

check_07

llms.txt: put it up, and expect nothing from it

There's a proposed standard, llms.txt, where you give language models a summary and a signpost in your site's root folder. The honest picture isn't a cheerful one: no major AI provider has officially committed to using it, and an analysis of 300,000 domains found no statistical link to getting cited. It takes half an hour and can't hurt, so put it up. But anyone who sells it to you as the solution to AI visibility is selling you hot air. The six checks above are the ones doing the actual work.

the_verdict

What your outcome tells you

Look back at where those checks actually live. A file in the root folder. A bot setting in a CDN dashboard. A step in your publishing process. HTML that's readable without JavaScript. Structured data in the source code. Not one of them is text you type into an editor.

If a few of them come up short, the question is rarely what needs to change. That's usually clear within the hour. The question is who does it. If your site sits behind a stack of plugins or with an agency that holds the keys, every single point becomes a ticket, a queue and an invoice.

And it doesn't stop after one round. In the past six months alone, one search engine added an AI report, another added an opt-out, new crawler names showed up, and your robots.txt became something your CDN can write to as well. The question isn't whether you can make one change, it's whether you can keep making them.

many checks come up short

Not one of these points is text you type into an editor. It lives in a file in the root folder, a CDN dashboard, your publishing process and your HTML.

a few checks come up short

What needs to change is usually clear within the hour. The question is who does it: behind a stack of plugins or with an agency that holds the keys, every point becomes a ticket, a queue and an invoice.

all good

It's right today. In six months an AI report, an opt-out and new crawler names showed up; the question is whether you can keep making changes.

no point score: the outcome mostly tells you who can make the changes, and whether that keeps working
now_what

Why we're so strict about this

We build websites that are fully in-house: your own code, your own hosting, nothing between you and your robots.txt. Not because in-house is inherently better, but because this kind of maintenance otherwise just doesn't get done. What that switch looks like is on the page from WordPress to in-house.

Want the background first? Read becoming visible in ChatGPT and Perplexity: it covers the five layers underneath this checklist. It's the same dependency you get with your business software, see Data Act: from 2027, switching can't cost you anything.

frequently_asked

Frequently asked questions about the AI visibility checklist

How do you check yourself whether your site is visible to AI assistants?

With seven checks you can run in an afternoon: read your robots.txt the way a crawler receives it, know which bot you are blocking and why, check whether you are in Bing and being cited there, check that your publishing process actually sends an IndexNow signal, search for a literal sentence in your page source, check that every important page has a valid schema block, and put up llms.txt without expecting anything from it.

Does blocking GPTBot remove me from ChatGPT answers?

No. Training, search and citation run through separate crawlers with separate consequences. Blocking GPTBot refuses the use of your content for training, but does not take you out of the answers. If you want to stay visible you may refuse training, but never search and citation. "Blocking AI" is not a single switch.

Why might my robots.txt differ from the file I manage myself?

Because intermediate layers can write to it. On 21 August 2026 Cloudflare announced Bot Preference Sync, which makes your robots.txt follow the AI bot policy from the Cloudflare dashboard; it rolls out from the week of 24 August across all plans and is on by default for new customers. So put the file as your browser receives it next to the file in your project, and know which layer accounts for the difference.

Where can I see whether an AI assistant is citing my pages?

In the AI Performance report in Bing Webmaster Tools, which Microsoft took out of preview on 11 February 2026. It shows how often your pages are cited in Copilot and on which questions. Google Search Console has reported separately on generative AI features since 3 June 2026, but only gives impressions.

What if my text only appears after a script runs?

Then AI crawlers risk seeing an empty page, because many of them run no JavaScript or only a limited amount. The check is simple: open the page source and search with ctrl-F for a literal sentence from your text. If you cannot find it, then as far as a crawler is concerned your text is not there.

Keep reading: the whitepaper

In the whitepaper, you'll read what the hidden cost of your current site really is, and how the switch works in one week.

Download the whitepaper