Writing · SEO & Content · Mar 27, 2026 · 7 min

AI search optimization (GEO): what actually matters now

Working notes on how AI answer engines actually pick their sources: unlinked mentions over backlinks, extractable structure over polished prose, and the robots.txt audit I ran on my own domains before telling you to run it on yours.

0.664 vs 0.218mentions vs backlinks
75,000 brandsin the correlation study

I've been reading a lot about how AI search engines (Google AI Overviews, ChatGPT answers, Perplexity) decide whose content to show. Ahrefs put numbers on it in 2025: a correlation study across 75,000 brands, plus a robots.txt crawl of roughly 140 million websites. I wanted to write down what stood out in the form I'd actually send a friend, then run their checks on my own site, because an untested summary of a study about earning citations is exactly the kind of page the study says gets ignored.

The short version: AI search rewards content that is easy to extract, answers early, and gets mentioned a lot. Rankings and crawlability still decide who gets found. Structure, freshness and topical trust now outweigh raw backlink volume.

Mentions beat backlinks

The most surprising finding: AI systems don't care about backlinks as much as they care whether people mention you, even without a link. In the 75K-brand data, brand web mentions correlate with AI Overview visibility at 0.664. Backlinks: 0.218. Three to one, roughly, in favor of people simply talking about you. Reddit threads, Quora answers, blog comments, niche communities. Two decades of SEO instinct says that's backwards. The data overrules the instinct.

The weighting shiftedAhrefs, 75,000 brands, 2025The old instincttwo decades of SEO habitBacklinks0.218 correlationHead termsCRM, shoesDepth firstanswer buriedRefreshout of ideasWhat AI search rewardsAI Overview visibilityWeb mentions0.664 correlationQuestionsseven words and longerAnswer firstdetail underneathRefreshwhen the facts moveCrawlability still decides discovery. Roughly 6% of sites block GPTBot in robots.txt, many of them by accident.
The same page graded by two rulebooks. Unlinked mentions outrank backlinks by roughly three to one, and the pages that get cited answer the question early instead of building up to it.

If your brand name keeps showing up around a topic, the model assumes you belong in the answer. The inverse is the uncomfortable part.

If no one talks about you, AI assumes you don't exist.
the finding that stung the most

AI triggers on questions

You almost never see an AI Overview for “shoes” or “CRM.” You absolutely see one for “how to structure a multi-region onboarding system for SMB teams” or “what small businesses need to pass a KYB review.” Seven-word questions and longer are where AI steps in.

So plan for the way people talk to a model. Problem-shaped queries trigger AI answers. Head terms mostly don't.

Structure is the extraction layer

AI likes structure. The pages that get cited basically do the model's job for it: pose the question, answer it directly in the first sentence, then explain the detail underneath. Tables, lists, clean headings, anything easy to lift.

The result reads flat. But if you think about how a model “reads,” it makes sense: it wants to grab the block that already sounds like a textbook answer. The pages hardest to cite are the ones that bury the answer under three paragraphs of setup.

Freshness stopped being optional

AI systems reward recent content, and the bar for “recent” is higher than a rewrite with two words changed. They want new facts, new numbers, new opinions, and they reward visible revisions: updated publish dates, refreshed data, even minor edits that show the page still reflects current reality.

I used to treat content refreshes as what you do when you run out of ideas. After reading this, I made refreshes a standing line in the content calendar I run: any post gets reopened when the facts underneath it move. This post is that line executing. It first went up in September 2023, and the version you're reading is my 2026 rewrite, with the study numbers and my own audit results added. New facts, new date.

Each platform eats a different diet

Per Ahrefs' source breakdowns, the platforms choose very differently, which killed any hope of a one-strategy-fits-all playbook:

  • Google AI, leans heavily on Reddit and YouTube
  • ChatGPT, favors big-name news organizations
  • Perplexity, favors niche experts and vertical blogs

My audience is fintech B2B, ops leads and compliance people asking precise, ugly questions. If I could fund only one bet, it's Perplexity: the person typing “what documents does a KYB review actually require” is already Perplexity's core user, and niche practitioner content is exactly its diet. Second bet: Reddit presence feeding Google AI. ChatGPT's newsroom preference is the one I can't earn my way into, so I don't spend on it.

google airedditchatgptnews orgsperplexityniche expertsmy audience asks precise, ugly questions
Three engines, three diets: an audience asking precise practitioner questions maps to Perplexity first, Reddit-fed Google AI second, and skips the newsroom-favoring ChatGPT entirely.

The robots.txt own-goal

One funny, and slightly terrifying, detail: some sites block AI crawlers in robots.txt without knowing it. In Ahrefs' crawl of ~140 million sites (May 2025), 5.89% block GPTBot, 5.74% block ClaudeBot, 5.61% block PerplexityBot. Imagine doing everything above correctly and then telling GPTBot “please ignore my entire website forever.”

Before prescribing the audit, I ran it on my own two properties. hi-daniel.com: three lines, allow everything, sitemap declared. Clean. creative.hi-daniel.com, where my live demos sit: no robots.txt at all. Nobody is blocked, but nothing is declared to crawlers either. No sitemap, no signals. Score: 2 domains checked, 0 own-goals, 1 missing file now on the fix list. Total cost: two curl commands.

hi-daniel.com: three lines, allow everything, sitemap declared. creative.hi-daniel.com: no robots.txt at all. Two curl commands to find out.
Five-minute audit: open yourdomain.com/robots.txt and check for disallow rules against GPTBot, ClaudeBot, or PerplexityBot. Roughly 6% of sites block them, many unintentionally. Cheapest fix in this entire post. Check subdomains too; that's where mine came up short.

The four-line checklist

Boiling the whole study down to what I'm actually operating on:

  • Be mentioned more often, communities count, links optional
  • Answer real questions people actually ask, early in the page
  • Make the structure easy for machines to read
  • Refresh content more often than the old SEO instinct says is necessary

If you run content anywhere: do the robots.txt check today, subdomains included. Put a refresh line in next quarter's calendar and name which posts it reopens. Then pick the one platform whose diet matches your audience and feed that one. Drop the other two. Start with the mentions. The backlink spreadsheet comes after.

Questions I keep getting

Full disclosure: this section is an extraction block, short, structured, textbook-cadence answers built for the exact machines described above. Dog food, eaten in public.

What is AI search optimization?

The practice of making content easier for systems like Google AI Overviews, ChatGPT, Perplexity, and Gemini to understand, extract, and cite. It combines clear structure, topical authority, and current information so a page can be used directly inside an AI-generated answer.

Does AI search still depend on traditional SEO?

Yes. Indexing, crawlability, topical relevance, and search visibility still decide whether a page is discovered at all. The difference is that AI search additionally rewards clean answer blocks, stronger entity signals, and content that's easy to quote or summarize.

What type of content gets cited most often?

Pages that answer a specific question quickly, use headings and lists that are easy to extract, stay current, and demonstrate authority through practical detail. A direct explanation gets reused far more often than a vague opinion piece.

Suggested posts