home / knowledge base / visible in ai
knowledge base

Getting Found in ChatGPT and Perplexity: What You Need Technically_

More and more people get an answer without ever clicking a search result. Being visible therefore means: being mentioned in the answer itself. Most articles on the topic focus on writing style. In practice it's mostly about technology, and so about whether you can actually get into your own site yourself.

the_core

Why this matters now, not next year

Gartner expects search volume through classic search engines to drop by a quarter this year, and organic search traffic to decline by more than half. Predictions, with all the uncertainty that comes with them. But the market is visibly preparing for it.

Since June 3, 2026, Google has been reporting separately on generative AI features in Search Console: you can see there how often your pages show up in an AI answer. The rollout started with a portion of UK sites and is going worldwide, with no set date. For now you only see impressions, no clicks, positions or search terms. At the same time, as of June 17, 2026 you got a toggle to keep your content out of those AI features entirely.

ai-search

which software suits an installation company with twenty engineers?

> 14 sources read, 3 cited

For an installation company of that size, three things matter most: scheduling engineers, job sheets on the phone and a link to accounting[1]. Separate packages usually cover one of the three; if you want them in one place, choose your own data layer with apps around it[1][2]. When choosing, look closely at what happens to your data if you ever want to switch[3].

  1. yourcompany.com/installation-sector
  2. a-comparison-site.com/software-for-installers
  3. a-trade-site.com/data-and-switching
the goal in one picture: not a link in a list, but the source under the answer; that is mostly settled in the technology of your site

That's the real news: not that AI search exists, but that it's becoming measurable and that choices now come attached to it. Measurability is usually the moment a hype turns into a discipline.

And what there is to do about it is rarely a list of writing tips. Almost everything that determines whether you end up in an AI answer sits in the technology of your site: in files, headers and settings. Five layers, from the bottom up.

five layers, from the bottom up; the bottom four do the work
layer_05

llms.txt: do it, but don't expect much

Half an hour of work and it can't hurt, but no major AI provider has committed to actually using it.

layer_04

Structure a machine can actually use

Structured data via schema.org and text that can be taken over without interpretation: question in the heading, answer in the first sentence.

layer_03

Text in the source code, not loaded into it

Many AI crawlers run little or no JavaScript. If your text only appears after a script, they see an empty page.

layer_02

Bing matters more than you'd expect

ChatGPT relies on Bing's index. Submit your site to Bing Webmaster Tools and turn on IndexNow.

layer_01

An AI crawler needs to be able to get in

Your robots.txt and firewall rules determine whether GPTBot, ClaudeBot, PerplexityBot and Google-Extended may fetch your pages.

the five layers of ai visibility: all changes to the site itself, not text you type into an editor
layer_01

First, an AI crawler needs to be able to get in

Before there's anything to cite, a crawler has to be able to fetch your page. Your robots.txt states who's allowed to. The names that matter: GPTBot, ClaudeBot, PerplexityBot and Google-Extended. The last one determines whether Google may use your content in AI answers, separate from your regular search results.

In practice those crawlers are surprisingly often set to block, without anyone consciously choosing that: a security plugin that stops "bots," a firewall rule at the host, a default setting that got switched on at some point. Cloudflare blocks mixed-use crawlers by default from pages with ads as of September 15, 2026, for new customers and all free accounts. Without ads that probably won't affect you. The point is where the switch lives: with your hosting, not in your text editor. So the honest question isn't whether you allow AI crawlers, but whether you know who manages your robots.txt and firewall rules.

layer_02

Bing matters more than you'd expect

By far the most traffic that comes to websites from AI assistants comes from ChatGPT: estimates put that share at around 78 percent. And for its web results, ChatGPT relies on Bing's index. For years, Bing was the search engine you could ignore, and that's exactly the one that now determines your visibility in the biggest AI channel.

Practically: submit your site to Bing Webmaster Tools and check whether your pages are actually indexed there. Also turn on IndexNow, the protocol that lets you ping Bing the moment you publish instead of waiting for a crawler. That requires a key file in your site's root directory and a call on every publish.

On a site you manage in-house, that's one line in your publishing process. On a site you don't manage yourself, it's a request to someone else.

layer_03

Your text needs to be in the source code, not loaded into it

Many AI crawlers run little or no JavaScript. If your text only appears after a script has fetched it, that kind of crawler sees an empty page. Google is more forgiving here than the rest, which makes it deceptive: you rank fine in Google and still show up nowhere in AI answers.

You can test it in ten seconds: right-click, choose "view page source" and search for a literal sentence from your own text. If you can't find it, an AI crawler probably can't either. Heavy page builders are the usual culprit here.

layer_04

Structure a machine can actually use

A language model prefers to cite what it can take over without interpretation. That means structured data via schema.org: Organization for your company, Article for your articles, FAQPage for frequently asked questions. It's a machine-readable explanation of what's on the page.

The shape of your text helps too. Put the question in the heading and the answer in the first sentence beneath it. Write out definitions in full: "an in-house data layer is ..." is citable, "here are the benefits" isn't. Mention numbers, dates and sources, because those are verifiable and get picked up more often. None of this comes with a guarantee: there's no ranking you can buy and no trick that pushes you into the answer. You're just making your page as easy as possible to use.

layer_05

llms.txt: do it, but don't expect much

There's a proposed standard, llms.txt: a small file in the root directory where you give language models a summary and a signpost. Honest picture: adoption sits at around two percent of sites, and no major AI provider has officially committed to actually using it. It takes half an hour and can't hurt, so we set it up. But anyone who sells it to you as the solution for AI visibility is selling you thin air. The four layers above are the ones doing the work.

the_test

Four things you can check yourself this afternoon

check_01

Open your robots.txt

Go to yourdomain.com/robots.txt and check whether GPTBot, PerplexityBot or Google-Extended are set to disallow.

check_02

Look at the page source

Open the source code of your most important page and search for a sentence of your own.

check_03

Count your pages in Bing

Search site:yourdomain.com in Bing and count how many pages show up.

check_04

Ask the AI itself

Ask ChatGPT and Perplexity who they recommend for your service in your region, and see who does get mentioned.

four checks for this afternoon, in this order: access, source code, index, answer

If one of the four falls short, the question is rarely what needs to change. That's usually clear within an hour. The question is who does it and how long you'll wait for it. Want to run through the full check? It's in the AI visibility checklist: seven checks, each with what a good outcome looks like.

now_what

Ultimately, this is a question of ownership

Look back at the five layers: a file in the root directory, a firewall rule, a key for IndexNow, structured data in the source code, HTML that's readable without JavaScript. Not one of these is text you type into an editor. They're all changes to the site itself.

If your website sits behind a stack of plugins, or with an agency that holds the keys, then each of those five layers is a ticket, a queue and an invoice. Manageable if it only has to happen once. But the rules of the game change every quarter: new crawler names, new reports, new standards that may or may not catch on. So the question isn't whether you can make one change, but whether you can keep making them.

That's why we advocate for a fully in-house site: your own code, your own hosting, nothing between you and your robots.txt. What that switch looks like is on the WordPress to in-house page. It's the same dependency you have with your business software, and the law is now shifting in your favor there too: see Data Act: from 2027, switching won't be allowed to cost anything.

frequently_asked

Frequently asked questions about visibility in AI search

How do you become visible in ChatGPT and Perplexity?

Mostly through technology, not writing style. Five layers decide it: an AI crawler has to be allowed to fetch your pages, your site has to be in Bing because ChatGPT leans on that index, your text has to be in the source code rather than appearing only after a script runs, your page needs structured data and quotable phrasing, and llms.txt is worth putting up without expecting much from it. The bottom four layers do the work.

Which AI crawlers should I allow?

GPTBot, ClaudeBot, PerplexityBot and Google-Extended are the names that matter. That last one decides whether Google may use your content in AI answers, separately from your regular search results. In practice these crawlers are blocked surprisingly often without anyone having chosen that: a security plugin stopping bots, a firewall rule at the host, or a default setting switched on long ago.

Why does Bing matter for visibility in ChatGPT?

Because for looking up current information ChatGPT leans on the Bing index. If you are not in Bing, you do not exist for that part of the search market, however well you rank in Google. Register your site in Bing Webmaster Tools and switch on IndexNow, so new pages do not have to wait for a crawl.

Can you measure whether you appear in AI answers?

Partly. Since 3 June 2026 Google reports separately on generative AI features in Search Console, but only gives impressions there: no clicks, positions or queries. Bing Webmaster Tools shows in its AI Performance report how often your pages are cited in Copilot, and on which questions. For now that is the fullest picture a search engine gives you itself.

Does schema.org help you get cited in AI answers?

It helps a machine understand what is on your page, but do not expect a breakthrough. Cited pages do contain JSON-LD more often, but that is a correlation: research on pages that added schema found no clear growth in citations. Treat it as hygiene, not leverage.

Further reading: the whitepaper

In the whitepaper about switching to in-house management, you'll learn what the hidden cost of your current site is and how the switch happens in one week.

Download the whitepaper