Periscopy
Articles
Methodology7 min read

Where the AI learns about your brand

A model does not have a database of companies. It has what it read. If your brand is not in the sources it read, it does not exist for the model, and absence raises no alarm.


Ask ChatGPT which is the best management tool for a Portuguese sales team. You get three names. Ask Perplexity the same and you get a similar list, with sources beside it. Ask Claude, and you get another variation on the same names.

In every case the answer came from somewhere. The model did not invent the brands. It learned them. And the question almost nobody asks is: learned them where?

The AI does not know your brand, it read about it

A language model does not have a database of companies. What it has is a statistical compression of an enormous amount of text it read during training and, in many cases, access to live web search at the moment of answering.

When you ask it for a recommendation it makes a synthesis: it pulls together what it saw about your space and returns the names that appear most often, in the most credible contexts, associated with the right words.

The corollary is uncomfortable: if your brand is not in the sources the model read, it does not exist for the model. It is not a penalty, it is absence. And absence raises no error, fires no alert and appears on no dashboard. It is silent.

The map of sources, by weight

Not all sources are worth the same. In rough order of influence over what a model says about a B2B brand:

1. Reddit, Hacker News, forums

Reddit is today one of the heaviest training sources for several models: OpenAI has a licensing deal and Google indexes it aggressively. Why? Because it is real human conversation: people recommending, complaining, comparing, with no marketing filter.

The models treat this as the closest thing to honest opinion they can find. One qualified mention in a relevant thread weighs more than ten pages on your own site. For Portuguese B2B: subreddits in your niche, Hacker News if you are technical, Quora, sector forums.

2. Wikipedia and Wikidata

These are the entity reference layer. They define what a company is, link it to a sector, to founders, to other entities. Being on Wikidata gives you a unique identifier that Google's knowledge graph and the models recognise.

Not being there means your brand is ambiguous: the model is not sure you exist as a distinct entity. The notability bar is real, and you need external sources that mention you. So it is not the first step; it is the consequence of having taken the others.

3. Review platforms and directories

G2, Capterra, Clutch, Trustpilot, Product Hunt. Ideal structure for a model: clear category, description, verified reviews, side-by-side comparisons. This is literally where the “best X” answers come from: the model read lists and rankings on these platforms and repeats them.

For a consultancy or an agency the equivalents are service-provider directories: Clutch, Sortlist, DesignRush, GoodFirms.

4. Publications and media

Articles, comparisons, interviews, “top 10 in [category]” lists. They give the model the narrative context: not only that you exist, but why you matter and what sets you apart. A mention in a credible outlet in your sector counts as third-party validation. It is exactly what is missing from whoever only has their own site.

5. Your own site

Necessary, but not sufficient. The model reads your site, and the better structured it is (schema, llms.txt, extractable content), the better it understands you. But it weighs it for what it is: an interested source. The site confirms and completes what the model already saw elsewhere. On its own, it rarely gets you into an answer.

The Portuguese asymmetry

Here is the pattern we see again and again in Portuguese B2B companies: 90% of the investment in visibility goes to point 5, the site itself, and almost nothing to points 1 through 4. The result is an impeccable site and an invisible brand inside ChatGPT. The work that can be controlled gets done and the work that matters gets ignored.

The good news is the reverse. In Portugal, points 1 to 4 are thinly populated in a B2B context: few Portuguese companies on Clutch, few qualified discussions on Reddit in your niche, few comparisons in Portuguese. Whoever goes in now occupies empty ground, an advantage that will not exist eighteen months from now.

What to do this quarter

You do not need all of it at once. In order of effort against return:

  1. Directories. Clutch, G2 Service Providers, Sortlist. An afternoon's work, with a consistent description across all of them. Approval in one or two weeks.
  2. Qualified presence on Reddit and Hacker News. Two or three threads in your space, with answers of real value and no obvious self-promotion. The models read the content, not the link.
  3. One article with data of your own. A benchmark, a study, numbers nobody else has. It is the way to generate external mentions without depending on media relationships.
  4. Wikidata, as soon as you have two external sources that mention you.

Your site stays as the solid base. The work that moves the needle is off-site.

How to know whether it is working

Define 50 to 100 questions real clients ask in your sector. Run them periodically in ChatGPT, Claude, Gemini and Perplexity. Measure how many you appear in (citation rate) and in what position relative to competitors (share of voice).

Without that baseline you are operating on faith. With it, it is management: you see the drift, you see the effect of each action, and you know exactly where you are still invisible.

In short

When ChatGPT recommends a company it is repeating what the internet says about that space, filtered through the sources it trusts most. GEO, in the end, is making sure the internet says the right thing about you, in the places where the model is going to read it.

See also: what entity authority is and how to measure citation rate and share of voice.