Blog · SEOAI search optimisation: getting cited by ChatGPT & AI Overviews
How answer engines pick their sources in 2026, what actually influences citation, and a concrete plan to make your site the one they quote.
A growing share of the questions your customers used to type into Google now get answered without anyone visiting a website. They ask ChatGPT. They read the AI Overview at the top of the results page. They ask Gemini, or Perplexity, or the assistant built into their phone. The answer arrives as a paragraph with two or three little citation links underneath.
If your business is one of those citations, you get the customer. If it is not, you are invisible in that conversation no matter how well you rank at position four on the traditional results page.
This has a name now: answer engine optimisation, or AEO. Most of what is written about it is either panic or snake oil. This guide is the practical version: how these systems actually choose sources, which of your existing SEO work carries over, which parts do not, and what to change on your site this quarter.
How answer engines choose what to cite
There are three broadly different machines behind the answers people see, and they behave differently. It helps to know which one you are optimising for.
1. Retrieval-augmented answers (ChatGPT search, Perplexity, Copilot)
The assistant reformulates the user question into several search queries, fires them at a search index (Bing powers a lot of this, and increasingly proprietary crawls), pulls the top handful of results, reads their text, and writes an answer from what it read. Citations are the pages it actually pulled text from.
The consequence is simple and under-appreciated: you cannot be cited if you were not retrieved, and you will not be retrieved if you do not rank for the reformulated query. Classic SEO is the entry ticket. What happens after retrieval is where the new work lies.
2. Google AI Overviews and AI Mode
Google decomposes a query into sub-questions, runs its own search for each, and synthesises. It draws heavily on pages that already rank in the top ten for the sub-questions, and it strongly favours pages with clearly delineated answers: a heading that matches the sub-question, followed by a direct answer in the first sentence or two.
Overviews cite more generously than people assume, and they cite pages that are not always the number one result. A page at position seven with a crisp, self-contained answer to a sub-question frequently beats the position one page whose answer is buried under 600 words of preamble.
3. Model memory, with no retrieval at all
When someone asks a model a general question with no live search, the answer comes from training data. You cannot optimise this directly and anyone who tells you they can guarantee it is selling something. What you can do is influence what gets written about you elsewhere: directories, industry publications, forums, comparison sites, review platforms. Models learn brands from the whole web, not from your homepage.
What carries over from ordinary SEO
More than the doom-posting suggests. If your technical SEO is sound, most of it works for you here too:
- Crawlability and indexation. Nothing gets cited that cannot be fetched and parsed.
- Server-rendered content. Some AI crawlers execute JavaScript; several do not, or do it inconsistently. Content that only appears after client-side hydration is a coin flip. Server rendering removes the gamble.
- Speed and stability. Crawlers have timeouts. A page that takes six seconds to return HTML gets abandoned.
- Topical authority. Sites with genuine depth on a subject get retrieved more often across the whole cluster of related questions.
- Links and mentions. Still the strongest signal of who is worth reading on a topic, and still the hardest thing to fake.
If you have not done the basics, do those first. AEO tactics layered on an unindexable site produce nothing. Our post on why your website is not ranking covers that groundwork.
What is genuinely different
Answer-first writing beats keyword-first writing
The single highest-leverage change is structural. Under every heading, answer the question the heading implies in the first two sentences. Then elaborate. Then give the caveats.
Traditional web copy does the opposite: it builds context, then reaches a conclusion at the end of the section. A language model summarising your page will happily lift a crisp opening claim. It is much less likely to reconstruct a conclusion you spread across eight paragraphs.
Practical rule: any section that a reader could reach directly from a search result should make sense on its own, with no dependence on what came before it.
Headings should be phrased as real questions or claims
“Pricing” is a poor heading. “What does DPDP compliance cost for a 20-person company?” is a good one, because it matches the sub-question an engine generates while decomposing a broader query. You are giving the retriever an exact lexical and semantic target.
Specificity is the citation currency
Models synthesising an answer prefer sources that contain concrete, checkable particulars: numbers, dates, named standards, thresholds, ranges, jurisdictions. Vague marketing prose contributes nothing to an answer, so it does not get quoted.
- Not “affordable pricing” but “₹65,000 to ₹2.4 lakh for a ten-page marketing site”.
- Not “fast turnaround” but “three to five weeks from kickoff to launch”.
- Not “we follow best practices” but “WCAG 2.2 AA, verified with axe and manual keyboard testing”.
This has a pleasant side effect: it makes your pages better for humans too, because specificity is also what a real buyer is scanning for.
Freshness signals matter more than they did
Answer engines are noticeably biased toward recently updated pages for anything with a temporal dimension: pricing, regulation, tooling, versions. Put a visible last-updated date on pages where recency is relevant, keep it honest, and actually revise the content when you touch the date. Systems that detect date manipulation exist, and the downside of being caught is worse than the upside of looking fresh.
Structured data helps parsing, not ranking
Schema.org markup does not make an engine like you. It makes your content unambiguous: this is an article, published then, by this organisation, about this. FAQPage, HowTo, Article, Organisation, Product and BreadcrumbList markup all reduce the guessing an extraction pipeline has to do. Reduced ambiguity means fewer reasons to skip you.
The pages that get cited most
Across the sites we run, the content types that pick up AI citations most reliably are:
- Definitional explainers that answer “what is X” without hedging, for a term with real search volume.
- Cost and pricing breakdowns with actual figures and the assumptions behind them. Almost nobody publishes real numbers, so the few who do get quoted constantly.
- Comparisons of two named things, where you state the conditions under which each one wins rather than declaring a universal winner.
- Checklists and step sequences with a stable number of steps and one idea per step.
- Regulatory and standards explainers that cite the clause, the deadline and the penalty.
- Original data: even a small survey of 80 customers, or a benchmark you ran yourself, becomes a source that others cite and models absorb.
Notice what is missing: company news, service pages full of adjectives, and thought-leadership essays with no verifiable content. Those can be excellent for brand. They are almost never cited.
Let the crawlers in, deliberately
AI crawlers identify themselves. You get to decide which ones you allow, and this is a commercial decision, not a technical one.
- GPTBot (OpenAI, training) and OAI-SearchBot (OpenAI, search indexing for citation).
- ClaudeBot and related Anthropic agents.
- Google-Extended, which governs Gemini training use without affecting normal Google Search indexing.
- PerplexityBot, Bingbot, Applebot-Extended and a long tail of others.
The distinction worth understanding: some of these bots feed training, some feed live retrieval and citation. Blocking a training bot costs you very little traffic. Blocking a retrieval bot removes you from answers entirely. A publisher protecting an archive may reasonably block training crawlers; a services business trying to be found should almost always allow the retrieval crawlers.
Check your robots.txt today. We regularly find sites that blocked everything with a wildcard during a staging deployment two years ago and never undid it.
Measuring it, honestly
Attribution here is genuinely poor, and anyone showing you a precise AI-traffic number is estimating. What you can actually measure:
- Referral traffic from chatgpt.com, perplexity.ai, copilot.microsoft.com and gemini.google.com in your analytics. Low volume, high intent, and it converts unusually well.
- Server log hits from the named crawlers above, which tells you whether you are even being fetched.
- Impression and click divergence in Search Console: impressions flat or rising while clicks fall is the classic signature of answers being served above your result.
- Manual citation checks: keep a list of 20 questions your buyers ask, run them monthly across the major assistants, and record whether you appear. Unglamorous, and the most reliable signal you will get.
A 90-day plan
If you want a sequence rather than a list of ideas:
- Weeks 1 to 2. Audit robots.txt and unblock retrieval crawlers. Confirm your key pages are server-rendered. Add Organisation and Article schema. Check server logs for AI crawler hits and fix anything returning errors.
- Weeks 3 to 4. Write down the 20 questions your customers actually ask before buying. Run each one through ChatGPT, Gemini and Perplexity. Record who gets cited. That list is your competitive set now, and it is often not who you think.
- Weeks 5 to 8. Restructure your ten highest-value pages: question-form headings, answer in the first two sentences, real numbers, visible update dates, self-contained sections.
- Weeks 9 to 12. Publish the two or three explainers that answer the questions where nobody in your market has a good page yet. Then start on the off-site footprint: directories, industry bodies, review platforms, credible mentions.
- Ongoing. Re-run the 20-question check monthly. Update pages where the facts have moved.
What not to do
- Do not spin up 200 AI-written pages. Answer engines are unusually good at ignoring generic content, and Google's spam policies target scaled content abuse directly.
- Do not hide instructions in your HTML aimed at models. It does not work, and being caught doing it is a credibility problem you do not need.
- Do not abandon traditional SEO. Retrieval runs on search infrastructure. Killing your rankings kills your citations.
- Do not chase every new AEO tool. Most of them wrap a prompt around the same three assistants and charge monthly for it. You can run the same checks by hand in an hour.
How RoseLeap can help
Our SEO & digital marketing work now assumes answer engines by default: server-rendered content, structured data, answer-first page architecture, crawler access reviews, and a monthly citation audit against the questions your buyers actually ask.
If you would rather start with a fixed-scope piece of work, the on-page SEO sprint covers the technical and structural foundation this all sits on. Tell us your market and your top ten queries on the contact page and we will come back with where you currently stand in AI answers, within one business day.
Rooted in Data · Built to Bloom
Need help with your own?
Tell us about your project. We come back with a clear, honest plan.