Periscopy
Articles
Guides7 min read

Six commands to find out whether the AI can read your site

Before asking why the AI does not cite you, confirm that it reaches the site. Six checks you run in the terminal in five minutes, with real output and what to do when each one fails.


The question we hear most is why the AI does not recommend the brand. Before answering that, it is worth answering a simpler one: can the AI even read the site? In around half the cases we audit, one of these six things is broken, and the content work done on top of it was being done blind.

Each command below was run against destaque.ai on 20 August 2026, and what is in the boxes is the real result, not an example drawn up for the article. Swap the domain for yours and it takes five minutes.

01 · The AI robots have a way in

The command

curl -s https://example.com/robots.txt | grep -iA1 'GPTBot\|ClaudeBot\|PerplexityBot\|Google-Extended\|OAI-SearchBot'

What it returned

User-Agent: GPTBot
Allow: /
--
User-Agent: OAI-SearchBot
Allow: /
--
User-Agent: ClaudeBot
Allow: /
--
User-Agent: Google-Extended
Allow: /
--
User-Agent: PerplexityBot
Allow: /

Five agents, five permissions. If any appears with Disallow, or does not appear at all while there is a generic Disallow above it, that engine does not read the site and nothing else matters.

Note the difference between the two kinds: GPTBot and ClaudeBot collect for training, OAI-SearchBot and PerplexityBot fetch the page at the moment somebody asks. Blocking the first kind is a legitimate decision; blocking the second is disappearing from answers with live search, which is where you win in the short term.

02 · The llms.txt exists and says something

The command

curl -s -o /dev/null -w '%{http_code} %{size_download}\n' https://example.com/llms.txt

What it returned

200  24759

Twenty-four thousand seven hundred and fifty-nine bytes of map in plain text (the 200 is the HTTP code): who we are, what we sell, which studies we published, with a line of context per link. A 404 here is not fatal, but it is an opportunity thrown away.

A ten-line llms.txt with only the page titles is no use at all. What makes it useful is the context sentence after each link, which is what an engine cites when it is not going to open the whole page.

03 · The text is in the HTML, not only in the JavaScript

The command

curl -s https://example.com/page | sed 's/<script.*<\/script>//g' | sed 's/<[^>]*>/ /g' | wc -w

What it returned

2082

Two thousand and eighty-two words arrive without running a line of JavaScript, on a page with one h1 and eleven h2. This is the test most sites fail, and they fail it silently: in the browser the page looks full, and what leaves the server is an empty shell.

If the number comes out near zero, the content is being assembled in the browser. Some engines run JavaScript, most collectors do not. The fix is to render on the server, not to add more schema on top of the emptiness.

04 · The schema declares what the page is

The command

curl -s https://example.com/page | grep -o '"@type":"[^"]*"' | sort | uniq -c | sort -rn

What it returned

14x  ImageObject
 1x  Organization
 1x  Person
 1x  WebSite
 1x  WebPage
 1x  Service
 1x  SoftwareApplication
 1x  FAQPage
 1x  BreadcrumbList

A product page that declares itself SoftwareApplication, with the organisation, the person who signs it, the frequently asked questions and the images identified one by one. Nine types, all of them true.

The common mistake is not having little schema, it is having schema that lies: declaring Product on a page that sells nothing, or FAQPage with questions that are not visible on the page. That gets caught, and costs more than having none.

05 · The versions in each language point at each other

The command

curl -s https://example.com/page | grep -oE '<link rel="(canonical|alternate)"[^>]*>'

What it returned

<link rel="canonical" href="https://www.destaque.ai/tracker"/>
<link rel="alternate" hrefLang="pt-PT" href="https://www.destaque.ai/tracker"/>
<link rel="alternate" hrefLang="en" href="https://www.destaque.ai/en/tracker"/>
<link rel="alternate" hrefLang="x-default" href="https://www.destaque.ai/tracker"/>

The Portuguese version declares the English one, and you have to run the same command on the English one to confirm it declares the Portuguese. Without reciprocity Google discards the whole pair and the two pages compete with each other.

This was a real mistake on that site: the homepage declared no alternative at all, and the English version was left with no signal on the most important page. Running the command on both sides is what catches it.

06 · The page exists in the sitemap

The command

curl -s https://example.com/sitemap.xml | grep -c '<loc>'

What it returned

64

Sixty-four addresses, and the pages that matter are in them, which is confirmed by searching the sitemap for each one. A sitemap that does not keep up with the new pages is worse than not existing: it gives the engine an out-of-date list and the engine trusts it.

Always check the pages created in the last few weeks. Those are the ones left out when the sitemap is written by hand instead of generated from the routes.

What these six commands do not say

They say the site is legible. They do not say the brand is cited, and the difference between those two things is the whole job. A site with all six points green may still not appear in a single answer, because being cited depends on authority, on original data and on answering the buyer's question better than the others do.

Legibility is the entry condition. Without it, nothing else counts; with it, the work begins.

Frequently asked questions

Do these six checks guarantee the AI will cite me?

No, and the difference matters. These checks say the site is legible: that the agents get in, that the text leaves the server, that the page declares itself for what it is. Being cited depends on something else, which is having authority and content that answers the question better than everyone else's. Legibility is the entry condition, not the victory. A site that is perfect on these six points may still not appear in a single answer.

Should I block GPTBot so the AI does not train on my content?

It is a legitimate decision and it depends on what you sell. What you should not do is block without distinguishing: GPTBot and ClaudeBot collect content for training, while OAI-SearchBot and PerplexityBot fetch the page at the moment a user asks. Blocking the second kind means disappearing from answers with live search, which is where visibility is won in weeks instead of months.

How often should these checks be repeated?

Once a quarter is enough for a stable site, and whenever you change framework, migrate server or launch a new section. The two that break by themselves most often are the third, when somebody converts a page to browser rendering, and the sixth, when pages are created that the sitemap does not catch.

Read next