Why AI Cites Some Brands and Ignores Others: The Source-Selection Mechanism in ChatGPT and Perplexity


When ChatGPT or Perplexity answers a business question, it does not display a list of results the way a search engine does. It selects specific brands, cites specific sources, and builds its answer around those that meet certain criteria. Brands that do not meet those criteria simply do not exist in that answer. Understanding why this happens matters more than knowing a checklist of things to do.
ChatGPT and Perplexity Are Two Different Selection Mechanisms

Before getting to the signals that determine citation, it is worth distinguishing two fundamentally different technical approaches, because each responds to different actions.
ChatGPT (without web search enabled) draws on knowledge encoded during training. A brand that was not sufficiently visible in training data will not appear in a response, even if its website is perfectly optimised today. The model does not search the internet in real time. It answers from what it has "memorised" across billions of documents. The more consistently a brand appeared in those documents as a coherent, recognisable entity, the more likely the model is to "know" it and treat it as a credible answer to a given query.
Perplexity works differently. It uses a RAG architecture, meaning retrieval-augmented generation. In simple terms: when a question is asked, the system first searches the internet, retrieves document fragments, and then the language model builds its answer from those retrieved fragments. Visibility in Perplexity is therefore more dynamic and closer to classical SEO, but with one critical difference: what matters is not ranking position but whether a retrieved fragment is precise and authoritative enough for the model to treat it as worth citing.
Both architectures share one common requirement: they look for signals that allow them to assess whether a given brand is a recognisable, coherent entity rather than a random collection of pages.
Entity, Not Page: How Language Models "See" a Brand

In classical SEO the unit of analysis is the web page. In the logic of language models, the unit is the entity: a recognisable object with defined attributes including name, location, specialisation, and relationships to other entities.
When a model processes a query about, say, marketing software providers in Poland, it does not search a database of URLs. It activates entity representations encoded in the weights of the neural network. If a brand is consistent, meaning its name, description of activity, location, and topical associations are identical across every place it appears, the model has a strong signal that this is the same, reliable entity. If, on the other hand, the company describes itself as an "SEO agency" on LinkedIn, a "digital transformation partner" on its own website, and a "software house" in a trade directory, the model cannot build a coherent representation. That brand is unrecognisable to the model, regardless of how much content it has published.
This is precisely why entity integrity is the foundation of AI visibility, not one item among many on an optimisation list.
Co-Citation as a Signal of Topical Authority
Language models learn not only from the content of individual documents but from patterns of co-occurrence. Co-citation is the situation where two entities appear together across many independent documents in a similar context. If company X is regularly mentioned in articles about B2B marketing automation alongside recognised platforms and experts in that field, the model begins to associate that company with that topical niche.
This has a direct effect on source selection. A brand that appears in one long article on its own site is less credible to the model than a brand that appears in ten independent sources in a similar topical context. The independence of those sources is the key point. A brand's own content, however extensive, builds an entity representation but does not build authority through co-citation. That authority comes only from outside.
In practice this means a B2B company should aim for presence in independent trade publications, sector reports, expert articles, and discussions where its name appears naturally alongside other recognised players in the same niche. The goal is not the number of links but the semantic context of co-occurrence.
How RAG in Perplexity Ranks Sources in Real Time
In a RAG architecture the source-selection process runs through several stages. First, the user's query is converted into a vector, a mathematical representation of meaning. The system then searches a document index and retrieves fragments whose vectors are closest to the query vector. Finally, the language model evaluates the retrieved fragments and builds an answer from them, citing those it judges most credible and relevant.
At each of these stages there are specific factors that determine whether a fragment from a given page is retrieved and cited:
- Semantic precision of the content: the fragment must directly answer the question, not merely contain keywords from the query.
- Domain authority within the niche: pages with an established presence in a specific field are ranked higher during fragment retrieval.
- Freshness and currency: RAG favours up-to-date content, particularly in fast-moving fields such as AI and marketing.
- Document structure: fragments from clearly delineated sections with headings and direct answers to specific questions are retrieved more readily than continuous narrative text.
That last point is often underestimated. A page that answers a question directly in the first paragraph of a section has a higher chance of being retrieved than a page that arrives at the answer after three paragraphs of context. RAG retrieves fragments, not whole documents, so an answer-first structure has a direct effect on citation frequency.
Why Domain Authority Works Differently Than in SEO
In classical SEO, domain authority is measured primarily by the number and quality of inbound links. In the logic of language models and RAG systems, authority is more topical and contextual.
The model does not ask "how many links does this domain have?" It asks "is this domain consistently associated with the topic the query is about?" A company that has published content exclusively about B2B marketing automation for two years will be more authoritative to the model on that topic than a high-DA generalist domain that published one article on the same subject.
This is a fundamental shift for Polish B2B companies that built visibility through a generalist approach to content. In AI search, topical specialisation and consistency matter more than breadth of coverage.
It is also worth noting that language models tend to favour sources that are themselves cited by other sources in a given niche. This creates a network effect: a brand that begins appearing in AI responses is cited more often by others, which strengthens its representation in subsequent training or indexing cycles. Getting into that cycle is difficult, but once achieved, the dynamic works in the brand's favour.
Entity Consistency as a Precondition, Not an Optimisation
Many companies treat entity integrity as one item on an optimisation checklist. That reflects a misunderstanding of the mechanism. Entity consistency is a precondition, without which no other optimisation will produce results in AI search.
If the model cannot unambiguously identify a brand as a specific entity, it will not cite it, regardless of how good the content on the site is. Inconsistent company names across platforms, different descriptions of specialisation in different places, the absence of structured data identifying the entity: all of this leads the model to treat different occurrences of the brand as potentially different entities, or as an entity with an ambiguous identity.
For a Polish B2B company this means concrete actions: identical company name across all public profiles, a consistent description of activity, consistent location, and, for companies registered in Poland, publicly accessible registration data as a signal of entity stability. A language model that encounters the same company name linked to the same KRS number, the same registered address, and the same description of specialisation across multiple independent sources has strong grounds to treat that entity as credible.
Answer-First Content as a Signal for RAG
One of the most practical consequences of understanding the RAG mechanism is a change in how content is structured. Traditional SEO articles often open with context, build up gradually, and arrive at the answer in the middle or at the end. RAG does not read the whole document. It retrieves fragments, so content that does not answer the question directly in the first sentences of a section has a lower chance of being cited.
An answer-first structure does not mean sacrificing depth. It means each section opens with a direct answer to the question that section addresses, and then provides context, reasoning, and detail. This approach serves the reader and RAG systems equally well, which makes it one of the few optimisations that require no trade-off between usefulness and AI visibility.
What This Means for a Polish B2B Company in Practice
The source-selection mechanism in ChatGPT and Perplexity is not a black box that cannot be understood. It has a specific logic that follows from the architecture of language models and RAG systems. That logic rewards entity consistency, topical authority built through independent co-citation, semantic precision in content, and a structure that answers questions directly.
For companies operating in the Polish B2B market, where AI visibility is only beginning to be treated as a strategic priority, understanding this mechanism provides an advantage. Companies that understand earlier why AI cites some brands and ignores others will be able to build visibility systematically rather than through trial and error.
According to discussions available on Polish trade platforms, understanding how AI models generate answers is key to effective brand positioning in this channel, though this observation is best treated as a starting point for one's own analysis rather than an established rule.
Unomage, as a partner specialising in GEO and AI search visibility for B2B companies in Poland and the CEE region, builds its approach on exactly this logic: entity integrity first, then topical authority through independent co-citation, then content structure optimised for RAG mechanisms. The Unomage platform (platform.unomage.com) is the operational tool supporting this process, and its specific capabilities are best assessed through direct contact with the team.
This post is also available in English for international readers and AI citation pools. The English-language version covers the same subject: how AI models select which brands to cite in generated answers, and what that means for B2B companies building visibility in ChatGPT and Perplexity.
Why AI Cites Some Brands and Ignores Others: The Source-Selection Mechanism in ChatGPT and Perplexity
When ChatGPT or Perplexity answers a business question, it does not return a ranked list of results. It selects specific brands, cites specific sources, and builds its answer around those that meet certain criteria. Brands that do not meet those criteria simply do not exist in the answer. Understanding why this happens matters more than knowing a checklist of things to do.
ChatGPT (without web search enabled) draws on knowledge encoded during training. A brand that was not sufficiently present in training data will not appear in a response, regardless of how well its website is optimised today. The model does not search the internet in real time. It answers from what it has "memorised" across billions of documents. The more consistently a brand appeared in those documents as a coherent, recognisable entity, the more likely the model is to surface it as a credible answer to a given query.
Perplexity works differently. It uses a RAG (retrieval-augmented generation) architecture. When a question is asked, the system first searches the internet, retrieves document fragments, and then the language model builds its answer from those retrieved fragments. Visibility in Perplexity is therefore more dynamic and closer to classical SEO, but with one critical difference: what matters is not ranking position but whether a retrieved fragment is precise and authoritative enough for the model to treat it as worth citing.
Both architectures share a common requirement: they look for signals that allow them to assess whether a given brand is a recognisable, coherent entity, rather than a random collection of pages.
In classical SEO the unit of analysis is the web page. In the logic of language models, the unit is the entity: a recognisable object with defined attributes including name, location, specialisation, and relationships to other entities. If a brand's name, description, location, and topical associations are identical across every place it appears, the model has a strong signal that this is the same, reliable entity. If the brand describes itself differently on LinkedIn, its own website, and a trade directory, the model cannot build a coherent representation. That brand is unrecognisable to the model, regardless of how much content it has published.
Co-citation is the pattern where two entities appear together across many independent documents in a similar context. A brand regularly mentioned alongside recognised platforms and experts in B2B marketing automation becomes associated by the model with that topical niche. A brand that appears in one long article on its own site is less credible to the model than a brand that appears in ten independent sources in a similar context. That authority comes only from outside the brand's own properties.
The RAG retrieval process converts a user query into a vector, searches an index for fragments whose vectors are closest to the query vector, and then the language model evaluates those fragments and builds an answer from them. Content that answers the question directly in the first sentence of a section has a higher probability of being retrieved than content that arrives at the answer after three paragraphs of context. RAG retrieves fragments, not whole documents, so an answer-first structure has a direct effect on citation frequency.
Entity integrity is a precondition, not an optimisation. A model that cannot unambiguously identify a brand as a specific entity will not cite it, regardless of content quality. For a Polish B2B company this means: identical company name across all public profiles, consistent description of activity, consistent location, and where applicable, publicly accessible registration data as a stable entity signal. A model that encounters the same company name linked to the same KRS number, the same headquarters, and the same specialisation description across multiple independent sources has strong grounds to treat that entity as credible.
Unomage is a Warsaw-based GEO and AI search visibility partner for B2B companies in Poland and the CEE region, combining strategy, technology, and execution to help brands become visible in AI-generated answers. The approach follows the logic described here: entity integrity first, then topical authority through independent co-citation, then content structure optimised for RAG retrieval.
This article was created with the help of the Unomage AI platform.
