Generative engine optimization: how to get cited by AI search
In short
- GEO (generative engine optimization) is the work of getting a page into the set of sources an AI answer links to. It sits on top of classic SEO: without indexing, speed and a clear structure there is nothing to cite.
- Format is the main lever. A direct answer in the first 200–300 characters, question-shaped headings, one idea per block, tables, lists, concrete numbers with dates and an FAQ. Models pull individual extractable fragments; whole articles never make it into an answer.
- The technical part takes about an hour to check: AI crawler access in robots.txt (there are three distinct classes, and blocking a training bot does not remove you from AI answers), Article, FAQPage, Person and Organization schema, and an author with a verifiable profile.
- As of August 2026 llms.txt is still a draft standard. Originality.ai tracked growth from 4,088 to 36,120 sites with the file over a year, while Ahrefs log analysis across 137,000 domains found that 97% of those files received zero requests in May 2026. No major vendor has confirmed it reads them.
- Attribution barely exists. Since June 2026 Search Console reports impressions in AI Overviews and AI Mode but no clicks, and AI Mode links carry noreferrer, so that traffic never shows up in analytics. Plan GEO with that limitation built in.
Someone searches for "how to calculate LTV in real estate", reads the answer right inside the results page and closes the tab. The site the answer was assembled from never gets the visit. Various 2026 measurements put AI Overviews on roughly half of Google queries; Yandex, the dominant search engine in Russia, reported a neural answer on about 42% of its searches. At Google I/O in May 2026 the company rebuilt the search box around AI Mode, keeping the link results alongside it.
For a website this reframes the task. A top-10 position still matters, since models assemble answers mostly from what sits in the index. But a second goal appears: being one of the sources the system quotes and links. The practice built around that goal is called GEO – generative engine optimization, sometimes AEO or LLMO.
Below is what works as of publication (August 2026), what you can verify on your own site in an evening, and where the honest answer is "there is no data". I run this blog on the same principles – schema, llms.txt, an FAQ in every article – so some examples come from it.
What changed in search by August 2026
Three shifts define how content works right now.
The answer pushed the link list below the fold. AI Overviews own the first screen. Click-through measurements differ in detail but agree on direction: when an AI block appears, organic CTR drops sharply, and the gap between cited and uncited sites grows wider than the drop itself. Being cited became an asset of its own.
Citations became specific. In the May 2026 update to AI Overviews and AI Mode, Google moved source links next to the individual statements they support, replacing the grouped list under the answer. The practical consequence: the unit of citation became a single paragraph or table row. Fragments are what you now optimize.
Russian-language search runs its own loop. In spring 2026 Alisa AI, Yandex's assistant, moved directly into the results page, and Yandex Webmaster (their equivalent of Search Console) added a section showing a site's visibility inside those answers. The selection logic is described in familiar terms: a direct answer near the top, structured lists and tables, a visible FAQ, an update date, an author byline, schema markup. Optimizing for Google and for Yandex barely diverges here.
GEO and classic SEO: what overlaps and what differs
GEO stands on the technical foundation of SEO. A page that is not indexed, loads slowly or hides its text behind scripts will reach neither the results page nor the model's answer. The difference lies in the target action and the unit of optimization.
| Dimension | Classic SEO | GEO |
|---|---|---|
| Goal | Page position in the results | Inclusion among the cited sources of an answer |
| Unit of optimization | A page for a query cluster | An extractable fragment: paragraph, list, table row |
| Query shape | Short keyword phrase | A full question, or several questions in a conversation |
| Key signal | Links, relevance, engagement | Clarity of phrasing, sourced facts, author authority, brand mentions |
| Measurability | Positions, impressions, clicks, conversions | Impressions in AI blocks, manual citation checks; clicks mostly invisible |
The classic SEO checklist stays fully in force: technical health, speed, internal linking, completeness against intent. GEO layers requirements about form on top. That is why GEO struggles as a standalone project and works well inside a content system where it is decided in advance who writes, about what, and to which template. I covered how that system is assembled in the article on systematic marketing.
Content structure an answer can be pulled from
This is the most controllable part of GEO. It is about form: the same material, arranged differently, gets cited differently.
A direct answer in the first 200–300 characters
The heading poses a question, the first paragraph answers it. No warm-up, no history of the industry, no promise to explain shortly. A model takes the first self-contained fragment that answers the user's question; if there is none, it assembles the answer from someone else's text. The test is simple: hide everything except the section's opening paragraph. If it reads as a finished answer, the format is right.
Question-shaped headings and one block per question
Write H2 and H3 the way a person would ask out loud: "How much does…", "What is the difference between…", "How to check…". Then keep it to one question, one block, one answer. If the answer is spread across three sections, it cannot be extracted whole, and the page loses to material that keeps it in one place.
Specifics instead of adjectives
There is nothing to quote in "pays back quickly". There is plenty to quote in "a pilot pays back in 4–6 weeks at a budget from $1,500 a month". A number, a unit, a period, a source. The date next to the number matters separately: models carry a sense of recency, and a claim anchored to "as of August 2026" holds up better than a timeless one. First-hand data – your own measurements, project breakdowns, results you can vouch for – outperforms rehashed reviews, because there is nowhere else to get it.
Tables, lists and FAQs
- Comparison tables are the most extractable format for "X or Y" and "what is the difference" queries. Column headers should make sense without the surrounding article.
- Numbered lists for sequences and instructions. Each item is phrased as a complete action that makes sense apart from the previous one.
- An FAQ block at the end with direct questions and 40–70 word answers. These are ready-made pieces for a model's answer and the basis for FAQPage markup.
- Definitions – a single sentence that explains a term without referring back to the rest of the text. Sentences like these are the ones that most often end up quoted.
There is a side effect: the same format helps a human reader who scans the page in 20 seconds. "Convenient for people" and "convenient for models" hardly diverge here. The tools I use to draft and check such fragments are listed in my roundup of AI tools for marketers.
Technical groundwork: crawlers, schema, llms.txt
AI crawler access in robots.txt
The first thing to check is whether the site lets in the bots that build AI answers. It helps to sort them into three classes, and confusing them is expensive.
| Class | Examples | What it does |
|---|---|---|
| Training | GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, Amazonbot | Collect data for training future models. No effect on today's citations |
| Search index | OAI-SearchBot, Claude-SearchBot, PerplexityBot | Build the index an answer is assembled from. Directly affect citation |
| User-triggered | ChatGPT-User, Claude-User, Perplexity-User | Open a page at the moment a live user in a chat asks for it |
A common mistake is blocking GPTBot out of "I do not want to be training data" and losing visibility along with it. GPTBot and OAI-SearchBot are governed by separate directives, and blocking the first does not block the second. Google is its own case: AI Overviews and AI Mode draw on the main index crawled by Googlebot, so disallowing Google-Extended does not remove a page from AI answers – that token covers model training and data grounding in Gemini products. The standard way to stay out of AI blocks is nosnippet and max-snippet, which kills your regular snippet too; a dedicated Search Console toggle went into limited testing in June 2026 for a subset of publishers.
On this site all three classes are allowed by an explicit list in robots.txt. The reasoning is simple: being read by models is a distribution channel. For a business that sells expertise, landing in a ChatGPT answer is worth more than shielding a paragraph from it.
Schema.org: the machine-readable subtitle track
Markup earns no citations on its own. Its job is to remove ambiguity: who the author is, when the page was updated, which part is a question and which is an answer. A minimum set for an expert blog:
- Article (or BlogPosting) – headline, description, datePublished and dateModified, author, publisher.
- FAQPage – the page's questions and answers in machine-readable form.
- Person for the author – with sameAs pointing to profiles where the identity checks out: site, social accounts, publications, talks.
- Organization for the company – logo, contacts and sameAs, so the brand resolves into a single entity.
- BreadcrumbList – the section structure, so the page reads in the context of its section and the site around it.
llms.txt: where it actually stands
The /llms.txt file was proposed as a sitemap for language models: a short markdown document listing key pages with brief explanations. Sensible idea, draft status, no official adoption. Originality.ai tracked growth from 4,088 sites with the file in June 2025 to 36,120 in May 2026. Ahrefs log analysis across 137,000 domains found that 97% of llms.txt files received zero requests in May 2026, and among those that did, SEO audit tools led the traffic rather than AI systems. Google has stated publicly that Search does not use llms.txt; OpenAI and Anthropic point webmasters to robots.txt instead.
Ship it or skip it
Ship it if it costs half an hour: the file is small, harmless, and may serve agents that browse your site on a user's behalf. Do not put it in the plan as a traffic source – as of August 2026 there is no confirmed evidence that llms.txt affects citation. This site has one, on exactly that reasoning and with no expectations.
E-E-A-T, authorship and mentions off your own site
Models pick sources that can be held accountable for a statement. That brings back signals familiar from E-E-A-T: experience, expertise, authoritativeness, trustworthiness. In practice it comes down to a few verifiable things.
- A real author with a track record. Name, photo, a bio made of checkable facts, an author page listing their work, Person markup, links to external profiles. An anonymous "editorial expert" produces none of that chain.
- First-hand experience in the text. Your own measurements, project details, constraints and failures. This is the layer no one can paraphrase from other articles, and it is the one that most often gets quoted.
- Transparent updates. A visible modification date and a short note on what was revised. For topics where data ages within a quarter, this is decisive.
- No contradictions between pages. If the homepage, the blog and the author profile claim different specializations and different experience, the brand entity resolves poorly.
Mentions beyond your own domain are a separate signal. Citation studies through 2026 converge on the finding that communities and reference sites take a large share of links in AI answers: Reddit, Wikipedia, industry directories, YouTube. At the same time the source mix differs sharply between systems – one measurement put domain overlap between ChatGPT and Perplexity at around 11%. The practical takeaway: no single platform guarantees anything, while presence in relevant discussions, directories, podcasts and guest pieces raises the odds of appearing in at least some answers.
The easiest way to verify results is by hand: ask 15–20 real questions from your audience in ChatGPT, Perplexity, AI Mode and Alisa, record who gets cited, and repeat monthly. I looked at how one of these systems handles its source interface in my Perplexity review.
What is measurable and what stays blind
Here comes the uncomfortable part that GEO articles tend to skip. Analytics for AI search is fragmentary as of August 2026.
Since June 2026 Google Search Console has reported generative surfaces separately – impressions in AI Overviews, AI Mode and AI features in Discover, broken down by page, country and date. There are no clicks in those reports. AI Mode links are served with a noreferrer attribute, so GA4 cannot recognize such a visit as coming from AI. GA4 added an AI Assistant channel group in May 2026 for referrals from ChatGPT, Perplexity and similar services, but clicks from Google's own AI blocks land in Organic Search, because they happen inside Search.
What this means for planning:
- 01Do not promise yourself or leadership a report on "traffic from AI". A complete picture does not exist; only indirect slices do.
- 02Track citations manually. A fixed list of questions, a monthly run across the systems, a table of who got cited. Cheap, and more honest than any estimate.
- 03Treat AI block impressions in Search Console and Alisa visibility in Yandex Webmaster as a direction indicator. They are not a sales metric.
- 04Add a "how did you hear about us" field to forms and first calls. Some answers of the "I asked ChatGPT" kind surface only this way.
- 05Judge GEO by the overall content result: branded queries, direct visits, the quality of inbound inquiries. If those grow, the work is landing even when the channel itself cannot be attributed.
What GEO does not give you
Nobody has a guarantee of citation. The source mix in AI answers shifts from query to query and from update to update; there is no bid and no setting that locks you into an answer. The workable frame is raising the probability of inclusion and accepting that part of the effect stays unmeasurable.
A sequence to follow if you are starting from zero: check robots.txt and schema, rewrite the opening paragraphs of key pages into direct answers, add FAQs and tables where they fit, put authorship in order. Two to three weeks of work, after which it becomes an ordinary content cycle with new format requirements. If you need a team to build that system and keep running it, that is what my agency Matveo does.
Frequently asked questions
What is GEO optimization in plain terms?
GEO (generative engine optimization) is preparing a site so its material ends up inside AI answers: AI Overviews and AI Mode in Google, Alisa AI in Yandex, ChatGPT, Perplexity. Where classic SEO aimed at a page position, GEO aims at a citation – a specific paragraph, list or table row the system links to in its answer. The technical foundation stays the same: indexability, speed, structure.
Does GEO replace traditional SEO?
No, GEO works on top of SEO. Models assemble answers mostly from pages already present in the search index, so indexing, speed, internal linking and relevance remain mandatory. GEO adds requirements about form: a direct answer at the start of a section, question-shaped headings, extractable fragments, schema markup and verified authorship. Classic organic results have not disappeared either – the link list is still there, just below the AI block.
Do I need an llms.txt file, and does it affect citation?
As of August 2026 there is no confirmed effect. Google has publicly stated that Search does not use llms.txt, and OpenAI and Anthropic refer webmasters to robots.txt. Ahrefs log analysis across 137,000 domains found 97% of llms.txt files received zero requests in May 2026. It is still worth creating: it takes half an hour, harms nothing, and can help AI agents browsing your site on a user's behalf. Just do not count on it as a traffic source.
How do I check whether AI systems cite my site?
There is no automatic metric, so two approaches are used. First, the generative surface reports in Google Search Console (since June 2026 they show impressions in AI Overviews and AI Mode, but not clicks) and the Alisa AI visibility section in Yandex Webmaster. Second, a manual run: build a list of 15–20 questions your audience actually asks, put them to ChatGPT, Perplexity, AI Mode and Alisa once a month, and record whose sources get pulled in. The manual method gives a fuller picture than any available analytics.
Ruslan Matveev
I build marketing as a system. Founder of Matveo, shipping AI products.
More on the topic