I've been reading a lot about how AI search engines (Google AI Overviews, ChatGPT answers, Perplexity) decide whose content to show. Ahrefs put numbers on it in 2025: a correlation study across 75,000 brands, plus a robots.txt crawl of roughly 140 million websites. I wanted to write down what stood out in the form I'd actually send a friend, then run their checks on my own site, because an untested summary of a study about earning citations is exactly the kind of page the study says gets ignored.
Three things changed how I write a page: answer in the first line, refresh when the numbers move, and chase mentions before links. Engines still have to crawl and rank you first.
Mentions beat backlinks
What surprised me most is that AI systems don't care about backlinks as much as they care whether people mention you, even without a link. In the 75K-brand data, brand web mentions correlate with AI Overview visibility at 0.664. Backlinks: 0.218. Three to one, roughly, in favor of people simply talking about you. Reddit threads, Quora answers, blog comments, niche communities. Two decades of SEO instinct says that's backwards. Big brands collect mentions and AI visibility for the same underlying reason, so the coefficient may be reading brand size more than a lever anyone can pull. It still changes what I'd fund first, because a small team can go earn mentions this quarter, and 0.218 is a thin return on link building.
If your brand name keeps showing up around a topic, the model assumes you belong in the answer. The inverse is the uncomfortable part.
If no one talks about you, AI assumes you don't exist.
AI triggers on questions
You almost never see an AI Overview for “shoes” or “CRM.” You absolutely see one for “how to structure a multi-region onboarding system for SMB teams” or “what small businesses need to pass a KYB review.” Seven-word questions and longer are where AI steps in.
So plan for the way people talk to a model. Problem-shaped queries trigger AI answers. Head terms mostly don't.
Structure lets a model lift you
AI likes structure. The pages that get cited do the model's job for it: pose the question, answer it directly in the first sentence, then explain the detail underneath. Tables, lists, clean headings, anything easy to lift.
The result reads flat. Watch how a model “reads” and it makes sense: it grabs the block that already sounds like a textbook answer. The pages hardest to cite are the ones that bury the answer under three paragraphs of setup.
Freshness stopped being optional
AI systems reward recent content, and the bar for “recent” is higher than a rewrite with two words changed. They want facts and numbers that have moved since the last version, and they reward visible revisions: updated publish dates, refreshed data, even minor edits that show the page still reflects current reality.
I used to treat content refreshes as what you do when you run out of ideas. After reading this, I made refreshes a standing line in the ThinkingAI editorial calendar I run for the newsletter and the blog: any post gets reopened when the facts underneath it move, and I reopen it with the pipeline. This post is that line executing. It first went up in September 2023, and the version you're reading is my 2026 rewrite, with the study numbers and my own audit results added.
Each platform eats a different diet
Per Ahrefs' source breakdowns, the platforms choose differently, which kills the one-strategy-fits-all playbook:
- Google AI, leans heavily on Reddit and YouTube
- ChatGPT, favors big-name news organizations
- Perplexity, favors niche experts and vertical blogs
My audience is fintech B2B, ops leads and compliance people asking precise, ugly questions. If I could fund only one bet, it's Perplexity: the person typing “what documents does a KYB review actually require” is already Perplexity's core user, and niche practitioner content is exactly its diet. I haven't placed that bet yet, so it's an untested call on my own audience. Reddit presence feeding Google AI comes second. ChatGPT's newsroom preference is the one I can't earn my way into, so I don't spend on it.
The robots.txt own-goal
Some sites block AI crawlers in robots.txt without knowing it, which is funny until it's yours. In Ahrefs' crawl of ~140 million sites (May 2025), 5.89% block GPTBot, 5.74% block ClaudeBot, 5.61% block PerplexityBot. Imagine doing everything above correctly and then telling GPTBot “please ignore my entire website forever.”
Before prescribing the audit, I ran it on my own two properties. hi-daniel.com: three lines, allow everything, sitemap declared. Clean. creative.hi-daniel.com, where my live demos sit: no robots.txt at all. Nobody is blocked, but nothing is declared to crawlers either. No sitemap, no signals. That is 2 domains checked, 0 own-goals, and 1 missing file now on the fix list. Two curl commands. What I do not have yet is an instrumented before and after on a property with a live content program.
The four-line checklist
Boiling the study down to what I'm actually operating on:
- Be mentioned more often, communities count, links optional
- Answer real questions people ask, early in the page
- Make the structure easy for machines to read
- Refresh content more often than the old SEO instinct says is necessary
Run the robots.txt check today, subdomains included; mine took two curl commands. Then put a refresh line in next quarter's calendar and name the posts it reopens, and feed the one engine whose diet matches your audience. The backlink spreadsheet can wait until those two are running.
Questions I keep getting
I wrote this section for the engines this post is about: short, structured answers in textbook cadence. Dog food, eaten in public.
What is AI search optimization?
You write so that Google AI Overviews, ChatGPT, Perplexity and Gemini can understand, extract and cite you. Those engines pull a page straight into an answer when it has clear structure, topical authority and current facts.
Does AI search still depend on traditional SEO?
Yes. An engine still has to crawl and index a page before it can quote it, which is why I checked robots.txt on my own domains first. On top of that, AI search rewards clean answer blocks, stronger entity signals, and passages you can quote unedited.
What type of content gets cited most often?
Pages that answer a specific question in the first line, then support it with headings and lists that extract cleanly, stay current and show authority through practical detail. Models reuse a direct answer far more often than a vague opinion piece.