Insights — Generative Engine Optimization

How AI Answer Engines
Choose Which Sources to Cite

An answer engine does not rank ten links and let you pick. It retrieves a few passages, writes one answer, and names the sources it leaned on. Everything that matters for visibility happens in that retrieval step — before any question of quality, brand or budget comes up. This is what we have watched it reward.

It is retrieval, not ranking

A search engine answers a query by ordering documents. An answer engine answers a question by assembling one. It turns the question into a set of searches, pulls back passages that look like they resolve it, discards most of them, writes prose from the ones it kept, and attributes what it used. The published answer is a summary of that selection.

The consequence is blunt and it is the reason most brands misdiagnose the problem. In classic search, being outranked is a matter of degree: you are on page two, and page two has traffic. In an assembled answer there is no page two. Either a passage of yours was in the set the engine kept, or your brand does not appear in the answer at all — and no amount of being nearly good enough changes that, because nothing about the near miss is rendered anywhere the buyer can see.

So the question worth asking is not "how do I rank for this". It is narrower and more answerable: what makes a passage survive that selection.

Every source it cited was an article

On August 24, 2026 we asked an answer engine which marketing agency helps pharma and healthcare brands get cited in AI search answers — a buying question, phrased the way a buyer phrases it. The answer named four firms and attributed eight sources. Every one of those eight was an article on a company's own publication: explainers, playbooks, research write-ups. Not one was a services page. This is our own observation on one query on one day, not an industry statistic, and we are reporting it as such.

The reason this is worth writing down is that it contradicts the intuitive fix. The intuitive fix, when a brand is invisible in AI answers, is to rewrite the services page: add the vocabulary the engine is using, tighten the copy, mark it up. We had already tested that on our own domain — a page carrying the exact phrases of the question, and the engine still did not retrieve it. The pattern in the citations explains why. The engine was not looking for a company that says it does the thing. It was looking for a passage that explains the thing, and then it credited whoever wrote the passage.

A services page and an article can contain the same claims and still behave completely differently here, because they are written for different readers. One is written for someone already deciding between vendors. The other is written for someone still trying to understand the problem — which is exactly the state a person is in when they ask an assistant a question instead of typing a keyword.

What a citable passage has in common

None of these are tricks and none of them are new. What is new is that they are now the difference between appearing and not appearing, rather than the difference between position three and position seven.

01

It answers without its neighbours

A passage gets lifted out of your page and dropped into someone else's answer. If it depends on the paragraph above it — "as we saw", "this is why", an unexplained "it" — it stops making sense the moment it is lifted, and a passage that stops making sense is a passage that does not get used. Write so that any single paragraph could be the only one a stranger ever reads.

02

It states the mechanism, not the outcome

"We deliver measurable AI visibility" is unusable: it asserts a result and explains nothing, so there is no question it answers. "An answer engine retrieves passages rather than ranking pages, which is why a page can contain the right words and still never appear" is usable, because someone asked that. Explanations get cited. Claims get skipped.

03

It is unambiguous about who published it

An engine that cannot resolve who is behind a page has a cheap reason to prefer one it can. A stable canonical URL, a named author, a publication date, and structured data that says the same thing the visible text says — none of this wins the citation on its own, but its absence is a reason to drop you at the point where the engine decides whether a source is worth attributing.

04

Someone else has mentioned it

This is the one that cannot be fixed by editing your own site, and it is usually the binding constraint. The firms that get named in these answers tend to be firms that appear in other people's text — directories, roundups, coverage, other companies' articles. If everything an engine can find about you was written by you, you are a single unsupported source, and it will reach for a claim it can corroborate.

The four properties, and how to tell whether a page has them
Property Present when Absent when
Self-contained Any single paragraph still makes sense read on its own, by someone who never saw the rest. It depends on the paragraph above it — "as we saw", "this is why", an unexplained "it".
Explains a mechanism It answers a question someone actually asked, by describing how something works. It asserts a result — "we deliver measurable AI visibility" — and explains nothing.
Resolvable provenance Stable canonical URL, a named author, a visible publication date, and structured data that says what the visible text says. The engine cannot tell who is behind the page, and has a cheap reason to prefer one where it can.
Corroborated elsewhere Directories, roundups, coverage or other companies' articles mention the brand independently. Everything an engine can find about you was written by you. You are a single unsupported source.

The fourth is the one that cannot be fixed by editing your own site, and it is usually the binding constraint.

How to tell whether any of it is working

The failure mode here is not doing the wrong work. It is doing work you cannot evaluate, for months, because the only metric you kept was whether you got cited — a metric that reads zero for a long time and then reads one, and tells you nothing in between.

Fix a small battery of questions phrased the way your buyer would phrase them, and run the same battery on a fixed cadence. Record the whole answer, not just whether you were in it: which sources were cited, what kind of page each one was, and whether the answer named companies at all or only directories. Those three things move before your citation does, and when they move they tell you which of the four properties above is still missing.

Record the zeros too, with their date and their method. A run of identical zero results is not a wasted month — it is what tells you the answer is stable and that the thing you changed was not the thing that mattered. That is a finding. Discarding it because it looks like nothing happened is how teams end up repeating the same fix.

One methodological note that is easy to get wrong: what you read in an assistant's answer is not a search ranking, and it should never be filed as one. They are different measurements of different systems, and mixing them produces a report that cannot be acted on.

Read: what fourteen days of recording our own zeros showed →

Where this comes from

The observation in section 02 is ours and is dated. Everything else that is not ours is below, so you can check it instead of taking our word for it.

Frequently asked

Is generative engine optimization just SEO with a new name?

No, because the unit of competition is different. Classic SEO competes for a position in a list of ten links, and a page that lands eleventh still exists. An answer engine does not publish a list: it retrieves a handful of passages, writes one answer, and cites the sources it used. A page that is not retrieved is not eleventh. It is absent. The work that follows from that difference — writing self-contained passages, publishing on a URL an engine can retrieve, earning mentions off your own domain — overlaps with SEO but is not the same list of tasks.

Why do service pages get cited less often than articles?

Because a service page answers the question "what do you sell" and an article answers the question the user actually asked. When an engine assembles an answer it is looking for a passage that resolves a question on its own, in the third person, without requiring the reader to already be shopping. Service pages are written in the second person and they close on a call to action, which makes them poor raw material for a neutral answer. That is a property of the writing, not of the URL: a service page that also explains the mechanism can be cited, and often is.

Does structured data make a page more likely to be cited?

Structured data does not persuade a model to pick you. What it does is remove ambiguity about what the page is, who wrote it and when, which matters at the retrieval step and matters again when an engine decides whether a source is worth attributing. Treat it as hygiene with a real cost of being wrong: FAQ markup whose answers do not appear in the visible text is a manual-action risk in Google, so the marked-up text and the on-page text have to be the same text.

Why can nobody give you a date for appearing in AI answers?

Honest answer: nobody who tells you a number can support it, because the interval depends on how often your pages are recrawled, how often the engine you care about refreshes its index, and whether anything outside your domain has started mentioning you. What can be committed to is the measurement, not the date: run the same battery of questions on a fixed cadence, record every result including the zeros, and you will see the shape of the answer change before the citation appears.

Can you optimize for ChatGPT specifically?

You can optimize for the retrieval layer that several assistants share, and that is most of the work. Beyond that, the differences are real but unstable: assistants change how they browse, which index they lean on, and how aggressively they cite, and they change it without notice. Building your content strategy around one assistant's current behaviour means rebuilding it the next time that behaviour changes. Build for the property they all reward — a passage that answers a question on its own, on a page that is unambiguous about who published it.

What is the first thing to fix if a brand is invisible in AI answers?

Find out whether the problem is retrieval or absence, because the fixes are opposite. Ask the engine the question your buyer would ask and read what it cites. If it cites competitors who say roughly what you say, your pages exist and are not being retrieved, and the bottleneck is usually off your domain — mentions, listings, anywhere your name appears in someone else's text. If it cites nothing in your category at all, the material simply is not written yet, and the fix is on your domain. Guessing which of the two you have is how budgets get spent on the wrong one.

Find out what the engines say about you

We run the battery, read the citations, and tell you which of the two problems you have — the one on your domain or the one off it. The audit is the same work described above, on your questions.