seo

Generative engine optimization: how to get cited by AI answers

Alex

Generative engine optimization: how to get cited by AI answers
The short version

Generative engine optimization (GEO) is the practice of writing pages that an AI answer engine can quote. The unit of success is not a rank, it is a sentence of yours reused inside an answer, with your domain named next to it.

Four things decide it: the engine can reach the page, it can lift a passage that stands on its own, it can check the facts in that passage against other sources, and you can tell whether any of it happened. The last one is where most GEO advice stops, and it is the part with the real traps.

What is generative engine optimization?

Generative engine optimization is the work of making a page quotable by a system that answers a question instead of listing links. It covers Google AI Overviews, ChatGPT search, Perplexity, Copilot and the assistants embedded in browsers. The page still has to be indexable and useful to a human. What changes is that a machine now reads it looking for a passage it can lift whole.

The name is recent, the mechanics are not mysterious. An answer engine retrieves a handful of candidate pages, extracts the passages that look like they answer the question, and composes a reply out of them. Your page competes at the level of the passage, not the document.

What does GEO change compared with classic SEO?

Almost nothing about the technical baseline, and almost everything about the writing. A page that is slow, blocked or thin loses in both games. A page that ranks well and buries its answer in the third paragraph of a five-paragraph section wins one game and loses the other.

QuestionClassic SEOGEO
What competes?The pageThe passage
What is the win?A position in a listA sentence reused, with your name attached
Who reads first?A crawler, then a humanA crawler, then a model, then maybe a human
What kills you?Being invisibleBeing unquotable: an answer that needs three other paragraphs to make sense
How do you measure?Impressions, position, clicksMentions in answers, plus the classic metrics, because there is no console for citations

The two are not in tension. Everything below also makes the page better for a human in a hurry, which is the honest reason to do it even if the AI traffic never arrives.

Why does a citation matter more than a rank on these surfaces?

Because on an answer surface there is often no list to be tenth in. The reply names two or three sources. You are in it or you are absent, and absent looks exactly like not existing.

The trade is real and worth stating plainly: a cited answer can satisfy the reader without a click. You gain presence and lose a visit. Whether that is a good bargain depends on what you sell. For a brand whose product is bought after research, being the named source in the research is worth more than the click. For a site monetised by page views, it is not.

How do you write an answer an engine can lift?

What makes a passage self-contained?

A self-contained passage answers its own heading in its first two sentences, without depending on anything above it. No "as we saw", no "this", no pronoun pointing at the previous section, no number whose unit was given three paragraphs earlier.

The test costs ten seconds. Copy a paragraph out of the page, paste it into a blank document, and read it. If it still answers a question on its own, an engine can lift it. If it becomes a fragment, the engine will either skip it or, worse, lift it and produce something that misrepresents you.

The practical shape: definition first, in 40 to 60 words, then the nuance. That order feels blunt to write and is the whole point. A model that scans the opening of a section and finds a complete answer has what it needs; a model that finds a warm-up paragraph moves on.

Which facts can an engine actually check?

A number with a unit, a date, a named entity, a price, a version. Those are the pieces an engine can cross-check against other pages, and cross-checking is what turns a plausible sentence into a citable one.

What it cannot check is an adjective. "Considerably faster", "the leading solution", "significantly better results" carry no verifiable content, which is why they are simultaneously bad for humans and useless for machines. Replacing one superlative with one measured number is the single highest-yield edit on most pages.

Two rules keep this honest. State the source of the number in the sentence that carries it, so a reader and a model can both trace it. And never invent a benchmark to fill a gap: a fabricated statistic that gets quoted is a reputational problem that outlives the page.

Should headings be written as questions?

Yes, when the section really answers one. A heading phrased as a question tells the retrieval step what the passage under it is for, and it matches the shape of what people actually type into an assistant, which is a sentence and not a keyword.

The failure mode is turning every heading into a question mechanically, including the ones that introduce a list or a definition. A heading that promises an answer and delivers a preamble is worse than a flat noun phrase, because it invites an extraction that then finds nothing.

How should the page be structured?

When is a table better than a paragraph?

Whenever the content is enumerable: a comparison, a set of options, a list of specifications, anything with two or more dimensions. Prose that describes a comparison forces every reader, human or machine, to rebuild the table in their head. A table hands it over already built.

The rule of thumb: if writing the paragraph makes you use the words "whereas", "on the other hand" or "respectively" more than once, it is a table.

What structured data is worth adding, and what is theatre?

Structured data does not make a page citable. It makes the entities on the page unambiguous, which is a different and smaller job. Add it where it describes something real: an Article with a genuine author and date, an Organization, a FAQPage whose questions are actually on the page.

One correction worth knowing, because it causes a lot of wasted effort: Google retired FAQ rich results in 2023. Marking up your FAQ no longer produces the expanding accordion in the search results, and it will not appear in the Rich Results Test. The markup remains valid and remains read by answer engines, which is the reason to keep it. Adding it expecting the old visual result is chasing something that no longer exists.

The anti-pattern is the reverse of theatre and is genuinely dangerous: declaring questions in a FAQPage that do not appear in the visible page. That is the class of markup that earns a manual action. Derive the markup from the page, never maintain it as a second copy that drifts.

Can the engine even reach the page?

This is the check people skip, and it is binary. If the crawler is blocked, everything above is irrelevant.

AI crawlers use their own user agents, and they are not covered by whatever you decided about Googlebot years ago. GPTBot, ClaudeBot, PerplexityBot, Google-Extended and several others each read robots.txt under their own name. If your file allows *, they are already allowed, and naming them explicitly anyway is a cheap, unambiguous signal.

A second convention, llms.txt, is a plain-text map of the site written for a model rather than for a crawler: the pages that matter, each with one line saying what it is. It is a proposal rather than a standard, and no engine promises to read it. It costs almost nothing and it forces a useful exercise, which is deciding in one sentence what each of your pages is actually for. Generate it from your own route map rather than maintaining it by hand, or it will be wrong within a month.

How do you measure whether any of this worked?

There is no Search Console for citations. Nobody sends you a report saying an assistant quoted you. What exists is a set of indirect signals, and a measurement tool that is more treacherous than its reputation suggests.

What Search Console can and cannot tell you

Search Console is the only free, first-party record of how a search surface treats your pages, and it sets four traps that quietly produce wrong readings. These are measured on a property we run, not repeated from a blog post.

The trapWhat actually happensWhat it does to your reading
A total is not the sum of the rowsGoogle anonymises rare queries. On one property, the reported total was 892 clicks while the 2,850 per-query rows summed to 560.Rebuilding a total by summing rows understated it by 37 %. Adding a second dimension makes the gap worse.
A domain property covers every subdomainA sc-domain: property mixes the marketing site, the app, the help centre and anything else on the domain. We measured five hosts in one property.A per-page report that does not filter by host blends unrelated sites with no error and no warning.
Aggregation flips silentlyAdding the page dimension switches aggregation from by-property to by-page. On the same window we saw 85,437 impressions summed by page against a 70,575 by-property total.Comparing a by-page figure to a by-property one produces a difference that is an artefact, not a change.
The days are Pacific daysA European job computing "three days ago" in local time can ask for a day Google has not closed.No error, simply fewer rows, which reads as a drop.

Two smaller ones matter for anyone comparing a before and an after. Finalised data lags roughly two days, so a baseline taken on same-day data is not reproducible and the comparison is meaningless. And average position is weighted by impressions: it does not add up across rows or days, and it only compares between two windows of the same width.

How do you know an engine has even re-read the page?

This is the check that separates a real result from a non-result, and almost nobody does it. If you rewrite a page and see no change two weeks later, there are two completely different explanations: the change did not work, or nothing has looked at it yet. They lead to opposite decisions.

Search Console's URL Inspection API answers it directly: it returns the last crawl time for a URL. Compare that instant to the moment you deployed. Two constraints to plan around: the quota is 2,000 inspections per day per site, and there is no batch form, so it is one call per URL. Inspect the pages you actually changed, not the sitemap.

How much this matters, measured on our own site hours after a deploy: the home page had been recrawled after the change went live, while its English counterpart had last been seen six days earlier, before it. Same site, same deploy, same day, two opposite verdicts. A report that treated both as "no impact" would have been wrong about one of them.

What are the usable signals for citations themselves?

Three, in decreasing reliability.

  • Referral traffic from the assistants. Some of them pass a referrer. It is a small, undercounted number, and it is real: a session attributed to an assistant is proof that a citation existed and was clicked.
  • Deliberate prompting. Keep a fixed list of ten to twenty questions your pages should answer, ask them on a schedule, record which sources come back. It is manual and biased by personalisation. It is also the only method that tells you who is being cited instead of you.
  • Crawler hits in server logs. Filtering the access log by the AI user agents tells you what is being read, and how often. It says nothing about citation, but a page no crawler has ever fetched cannot be cited.

Set the expectation before you start: on a small non-brand cohort, click counts are too sparse to be significant. Position and impressions move first and are the readable metrics. Concluding from three clicks is how a measurement programme discredits itself.

Where does an automated content tool fit?

Most of the work above is editorial discipline: an answer in the first two sentences, a number instead of an adjective, a table instead of a comparison written out. A tool helps with the part that is repetitive, which is producing that shape consistently, at a pace a person cannot hold.

That is what Wisewand is built to do. It has two modes. In Autopilot, it picks and schedules the topics of a rolling 30-day editorial calendar itself, writes about one article a day and publishes it to WordPress, Shopify, WooCommerce, PrestaShop or a webhook, at 77 EUR per month after a 3-day trial for 1 EUR. In Manual mode you drive: one credit is worth roughly one article, credits stack up and never expire.

What it will not do, and what nobody can honestly promise: guarantee a citation. Answer engines choose their sources, they change how they choose, and no supplier controls that. What is controllable is producing pages built to be quotable and then measuring what happens, which is the whole content of this article. If your catalogue is on Shopify, the Shopify side follows the same logic applied to products and collections.

FAQ

What is generative engine optimization (GEO)?

Generative engine optimization is the practice of writing pages that an AI answer engine can quote inside a generated reply. It targets Google AI Overviews, ChatGPT, Perplexity and similar surfaces. The unit that competes is the passage, not the page, so the work concentrates on self-contained answers and checkable facts.

How does GEO differ from traditional SEO?

The technical baseline is the same: a page that is blocked, slow or thin fails at both. The difference is the target. SEO aims at a position in a list of links; GEO aims at a sentence of yours being reused, with your domain named. In practice that means answering the heading in the first two sentences rather than building up to a conclusion.

Does GEO replace SEO?

No. An answer engine still has to find and crawl the page, and it usually finds it the same way a search engine does. GEO is an additional layer of editorial discipline on top of a healthy technical baseline, not a replacement for it. Everything that makes a page citable also makes it clearer for a human reader.

What makes a passage "self-contained"?

It answers its own heading without depending on anything above it: no unresolved pronouns, no "as we saw", no number whose unit was defined earlier. The test is to copy the paragraph into a blank document and read it. If it still answers a question on its own, an engine can lift it safely.

How can I make my facts verifiable by an engine?

Use numbers with units, dates, versions and named entities, and state where each one comes from in the sentence that carries it. Those are the elements an engine can cross-check against other sources. An adjective such as "considerably faster" carries nothing checkable, which is why it helps neither a reader nor a model.

Should I turn every heading into a question?

Only where the section genuinely answers one. A question heading signals what the passage under it is for and matches how people phrase requests to an assistant. A heading that promises an answer and then delivers a preamble is worse than a plain noun phrase, because it invites an extraction that finds nothing.

When should I use a table instead of prose?

Whenever the content is enumerable: comparisons, option sets, specifications, anything with two or more dimensions. Prose describing a comparison forces every reader to rebuild the table mentally. A practical signal: if the paragraph needs "whereas" or "on the other hand" more than once, it should be a table.

Does structured data get my page cited?

No. Structured data removes ambiguity about the entities on a page, which is a smaller and different job. Worth knowing: Google retired FAQ rich results in 2023, so FAQPage markup no longer produces the search-results accordion and will not show in the Rich Results Test. It stays valid and stays read by answer engines, which is the reason to keep it.

Do I need to allow AI crawlers explicitly in robots.txt?

If your file allows *, they are already allowed. Naming GPTBot, ClaudeBot, PerplexityBot and Google-Extended explicitly costs nothing and removes any ambiguity. The check is worth doing because these agents are not covered by rules written for Googlebot, and a blocked crawler makes every other optimisation irrelevant.

How can I track whether an AI is citing my content?

There is no console for citations. Three usable signals, in decreasing reliability: referral traffic from the assistants that pass a referrer, deliberate prompting of a fixed question list on a schedule, and AI crawler hits in server logs. The last one proves the page was read, not that it was cited.

How do I know whether a change has even been seen?

Use Search Console's URL Inspection API, which returns the last crawl time for a URL, and compare that instant to your deploy time. Without it, "no impact" and "not looked at yet" are indistinguishable, and they call for opposite decisions. The quota is 2,000 inspections per day per site and there is no batch form, so inspect only the pages you changed.

How long before GEO work shows a result?

Plan on weeks, and check the crawl date before drawing any conclusion. Also set the expectation on what is readable: on a small non-brand cohort, click counts are too sparse to be significant, so position and impressions are the metrics that move first. Concluding from a handful of clicks is the fastest way to misread the whole exercise.

This article also exists in French: le lire en français