I gave the same prompt to five AI engines on May 10, 2026:
"Who are the best experts on AI SEO and AI automation for SEO right now? Give me names with brief context for each."
Five different answers. Some overlap. Some surprising omissions. And one engine that's making things up so confidently you'd believe it.
This isn't a leaderboard post. It's a data-journalism look at how AI engines decide who to cite, because the patterns matter more than the names. If you're trying to be cited yourself, the patterns are what you optimize for.
Why this is worth measuring
More and more, the first thing someone does when they need an expert is ask a model instead of running a search. "Who should I follow on this," "who actually knows about that," "who should we hire." Whatever names come back get the attention, and everyone else stays invisible. That makes being cited a distribution problem rather than a vanity one, and it's a problem almost nobody can see, because there's no Search Console for it. You cannot open a dashboard and find out how often ChatGPT names you.
So the field runs on theory. There's no shortage of advice about how to get cited by AI, and very little of it comes with a test attached. The gap between "here's what I reckon works" and "here's what happened when I checked" is where this study lives.
The re-run exists for a blunter reason. A study nobody repeats can't be wrong. It sits there being quoted, aging quietly, with no way for anyone to find out whether it was ever true. Repeating it is the only thing that separates a finding from a screenshot, and repeating it is how I found out that most of my own May results were noise.
What this study found
- Four names came back in 7 to 8 of 9 runs: Mike King, Aleyda Solis, Lily Ray and Kevin Indig. All four publish a long-running newsletter under their own name.
- Twelve names appeared in exactly one run out of nine. Single-run citation studies report those as findings.
- Follower count showed no correlation with citation. One of the least-cited names has over 50,000 followers.
- The engine that read eight sources and filtered the gamed ones gave stable answers. The engine that read two listicles gave a different answer every time.
- One model read an agency's "top experts" article and then recommended that agency's founder.
The five engines tested
- ChatGPT (GPT-5, web browsing on) — accessed via chatgpt.com
- Claude (Sonnet 4.6) — Claude.ai, web search enabled
- Perplexity (Pro, Sonar Large default model)
- Google Gemini (2.0 Pro)
- Bing Copilot (Balanced mode, web grounding on)
Same prompt. Same day. Each engine asked from a fresh session with no chat history.
The aggregated results
I asked each engine for ~8 names. Here's the merged frequency:
| Name | Cited by # engines | Notes |
|---|---|---|
| Aleyda Solis | 5/5 | Universal. Every engine. Strong personal brand + active newsletter + frequent speaking. |
| Kevin Indig | 5/5 | Universal. Growth Memo newsletter is heavily cited across engines. |
| iPullRank / Mike King | 4/5 | Missed only by Gemini. Lily Ray-adjacent crowd, strong technical citations. |
| Lily Ray | 4/5 | Missed only by Perplexity. Strong YMYL + algorithm-update voice. |
| Eli Schwartz | 3/5 | ChatGPT, Claude, Bing. Product-led SEO framing. |
| Ross Hudgens (Siege Media) | 3/5 | ChatGPT, Perplexity, Bing. Agency-affiliated but cited as individual. |
| Marie Haynes | 2/5 | Claude + Gemini. Strong on Google updates specifically. |
| Britney Muller | 2/5 | ChatGPT + Claude. AI-SEO crossover work. |
| Rand Fishkin | 1/5 | Only Gemini. Interesting — once-universal, now mostly fading from AI citations. |
| Glenn Gabe | 1/5 | Only ChatGPT. Strong technical + algorithm voice. |
And then there's the engine-specific tail: 14 different names cited by exactly one engine, including 3 that I'm fairly sure don't exist as SEO experts at all.
The Gemini hallucination
Gemini gave me "Maria Henderson, a thought leader at Search Influence specializing in semantic SEO and AI integration." Maria Henderson is not a notable SEO figure. Search Influence is a real agency. The combination is fabricated.
This isn't unusual — every engine I tested hallucinated at least one name. But Gemini was the worst, fabricating 2 of its 7 names. ChatGPT fabricated 0. Claude fabricated 0. Perplexity fabricated 1 (gave an outdated affiliation for a real person, said they were at Brafton when they left in 2022). Bing fabricated 1.
The takeaway: AI engines are not yet reliable for entity-level questions. Trust the patterns, distrust the individual citations.
What the patterns reveal
Pattern 1: Newsletter-driven brands dominate
The top 4 names (Solis, Indig, King, Ray) all have active, well-distributed newsletters. The single most predictive factor for cross-engine citation appears to be having a personal newsletter with a consistent multi-year publication history.
Why this works: newsletters are crawled and indexed. They get linked to from other newsletters. They build clear entity associations between a name and a topic over time. AI engines see the pattern across multiple sources and assign higher confidence.
Pattern 2: Twitter/X presence matters less than you'd think
I cross-checked each name against their X follower count. Zero correlation with citation frequency. Two of the top 5 cited names (Indig, King) have moderate X presence; two (Solis, Ray) have huge X presence. One of the bottom 5 (Britney Muller) has 50K+ followers.
The signal that does correlate: how often the person is quoted or referenced in third-party industry blogs (SEJ, SEL, Backlinko, Wynter). That citation density appears more important than direct social presence.
Pattern 3: Agency-affiliated experts get cited less than expected
Several well-known agency founders got 0/5 citation. The pattern seems to be: AI engines associate the brand with the agency, not the founder. Where the founder maintains a strong personal-brand newsletter or speaks under their own name, the citation lifts.
Pattern 4: Recency matters, but not how you'd think
I expected Perplexity (which heavily weights recency) to favor newer voices. It didn't. Perplexity cited the same established names. The recency advantage shows up in topic-specific citation — when I asked about "AI Overviews optimization specifically," Perplexity surfaced more recent voices than ChatGPT did.
So: for entity-level questions, recency doesn't help. For topic-specific questions, recency starts to matter.
The interesting absences
People I'd expect to see who got 0/5 citations:
- Andrew Holland (HelloBlue) — well-known in AI-SEO circles, no AI engine surfaced him
- Sam Torres (The Gray Dot Co) — same
- Maxim Salnikov — Microsoft-side AI SEO voice, surprisingly absent
- Dan Petrovic (Dejan SEO) — technical-SEO niche, didn't appear
None of these are obscure. Their absence likely reflects citation graph effects more than ranking effects — they're cited within SEO communities but less in cross-niche content that AI engines see.
The opportunity for new entrants
If you're trying to be cited as an AI SEO expert by these engines, the data suggests four moves:
- Newsletter, not Twitter. A 2-year-old newsletter beats a 200K-follower account on X.
- Get quoted in third-party publications. SEJ, SEL, Backlinko, Wynter, MarketingProfs. One quote in each per year does more than 10 LinkedIn posts.
- Use your own name, not your agency's. If your byline says "by Company X," engines associate the work with the company. Publish as yourself.
- Cover topic-specific subniches. Recency wins on specific topics. "AI Overviews," "llms.txt," "GEO" — easier to anchor a citation than "AI SEO" broadly.
The method (for anyone who wants to replicate)
If you want to reproduce this on your own niche (replace "AI SEO" with whatever), the setup is:
- Identify the 5 main AI engines (the ones I used are reasonable; you can add Brave Leo, Mistral Le Chat, etc.)
- Pick a question that's specific to your niche, generic enough to invite many names
- Run the same prompt, same day, from clean sessions
- Record the answers
- Cross-check fabrication — Google each name's affiliation to verify
- Tabulate frequency, count overlaps, look for missing-but-expected names
The total time investment: about 90 minutes for the data collection, another 90 for the writeup.
One pitfall: each engine remembers prior queries within a session. Run each prompt in a fresh chat. For Perplexity specifically, also use private/incognito browsing — it personalizes based on past queries on the same device.
Where I land
AI citation is patterned. The patterns are reproducible. The patterns favor:
- Multi-year newsletters with consistent publication
- Cross-publication quotes (not just self-published)
- Personal brand over agency brand
- Topic-specific authority on emerging subniches
None of those are quick wins. But none require huge follower counts, paid placements, or agency-level resources either. They reward the kind of patient, compounding work that good SEO has always rewarded — just with a different audience watching.
Pun intended.
Update, August 2026: I ran it again, and then I ran it nine times
The study above is from May 10. Three months is long enough in this field that the results could have been quietly wrong for weeks without anyone noticing, including me, so on August 4 I went back to check. I ended up throwing out my own conclusion twice before the data settled. The version you're about to read is the third one. I'm leaving the wrong turns in, because they're the whole point.
The short version: the four names at the top held up. Almost everything below them was noise, and I only found that out because I stopped asking each engine once.
What changed in the method
In May I asked five engines one question each, wrote down the answers, and treated those answers as readings. That's what everyone does with this kind of study. It's also, as it turns out, close to useless below the top of the list.
This time I could only reach three engines, but I asked each of them the same question three separate times, in fresh sessions. Nine runs instead of five. Fewer engines, far more information, because the second and third runs tell you something the first one structurally cannot: whether the answer is a measurement or a coin flip.
The three engines, and their surfaces, matter for what follows:
- GPT, through the CLI setup I use for multi-model second opinions, with no web access at all. Pure recall, answering from what it absorbed in training.
- Claude, through the Claude Code CLI, which searched the web and cited its sources.
- DeepSeek V4 Pro, through opencode with direct provider auth, which also searched.
Gemini's individual CLI tier has been discontinued. Perplexity and Bing Copilot have no CLI at all. Those three are exactly the engines that fabricated in May, which becomes important later.
The results, as hit-rates
Instead of "cited / not cited," every name now has a score out of nine. Anything appearing in fewer than two runs is listed at the bottom, because that's where it belongs.
| Name | GPT (recall) | Claude (search) | DeepSeek (search) | Total |
|---|---|---|---|---|
| Mike King (iPullRank) | 3/3 | 3/3 | 2/3 | 8/9 |
| Aleyda Solis (Orainti) | 3/3 | 3/3 | 2/3 | 8/9 |
| Lily Ray (Amsive) | 3/3 | 3/3 | 1/3 | 7/9 |
| Kevin Indig (Growth Memo) | 3/3 | 3/3 | 1/3 | 7/9 |
| Cindy Krum (MobileMoxie) | 3/3 | 0/3 | 3/3 | 6/9 |
| Jason Barnard (Kalicube) | 1/3 | 2/3 | 1/3 | 4/9 |
| Wil Reynolds (Seer Interactive) | 3/3 | 0/3 | 0/3 | 3/9 |
| Rand Fishkin, Dan Petrovic, Marie Haynes, Szymon Słowik, Britney Muller, Ethan Smith, Bartosz Góralewicz, Andrea Volpini | appeared in 2 of 9 runs each | 2/9 | ||
| Twelve further names | appeared in exactly one run each | 1/9 | ||
The top four survived a much harder test
Mike King and Aleyda Solis came back in eight of nine runs. Lily Ray and Kevin Indig in seven. Across three engines, two different retrieval modes, and three separate askings each.
That's the same four names that led the May study, and it's now a considerably stronger claim than it was in May, because in May it rested on one question per engine. A name that survives repeated asking across models that don't share training data or retrieval isn't riding on whatever ranked this morning. It's in the models' picture of the field.
All four still have the thing the original study identified: a long-running newsletter published under their own name. Three months and a much more paranoid method later, that's still the pattern.
Twelve names appeared exactly once
Here's the part that should bother anyone who has read a citation study, including mine.
Across nine runs, twelve different people were named exactly once and never again. Not obscure people, in most cases. Just names that surfaced when a particular engine happened to read a particular article on a particular attempt.
A single-run study reports every one of those twelve as a finding. It writes them into a table next to Aleyda Solis with no way to tell them apart. My May study did that. The first version of this August update did it too.
If you take one thing from this piece, take that: below the top few names, a single-run citation study is mostly reporting which listicle got retrieved that minute.
I got the reason wrong, twice
My first theory was that searching the web is what makes answers unstable, and that recall would be rock solid. It fit the first data I had: GPT, with no web access, returned six of the same eight names in all three runs, while DeepSeek returned one.
Then Claude's runs came in. Claude searches the web, cites its sources, and returned the top four in three out of three. So much for that theory.
Before that, I'd made the opposite mistake in the other direction. Having seen DeepSeek scatter, I concluded the top-four finding was broken and wrote a correction saying so. That correction was based on one engine. Which is precisely the error this whole exercise exists to catch, made by me, in the middle of writing about it.
Two reversals in one afternoon, both from reading too few samples.
I'm keeping that in because a study about how confidently machines state unreliable things shouldn't quietly launder its own author's confident unreliable statements.
What actually drives the instability: how many sources, and whether the model filters them
The difference isn't search versus recall. It's what a model does once it searches.
DeepSeek read two listicles and largely repeated them. On one run it read an article on primaryposition.com and then recommended David Quaid of Primary Position, along with a cluster of link-building names from that same page. The model read an agency's list, and recommended the agency.
Claude read eight sources and openly reasoned about which of them were rigged. Its exact words: the tell for a gamed list is that "the author's own agency appears in it, the entries have no specific claim attached, and the same eight names appear in the same order across five sites." Then it discounted those and picked people who had published a method or a number you could replicate.
Same question, same afternoon, two searching models. One got played by the citation supply chain. The other described the mechanism out loud and routed around it.

That is the practical finding for anyone trying to get cited. Being in the listicles will get you named by the models that skim. Only publishing something specific and testable under your own name gets you named by the models that read properly, and those are the ones whose answers don't change between Tuesday and Wednesday.
The riser nobody would have predicted
Cindy Krum of MobileMoxie scored 6/9. She was 3/3 on GPT and 3/3 on DeepSeek, and 0/3 on Claude. She does not appear anywhere in the May results table.
She's also the only person in the study who is perfectly stable on two engines while being completely absent from a third, which is a strange and specific pattern I don't have a confident explanation for. Entity-first, mobile-first SEO has been her subject for well over a decade, which fits the "narrow topic, held for years" pattern. Beyond that I'd be guessing, and I've done enough of that today.
The hallucination question, still unanswered
I checked every name and affiliation in all nine runs against primary sources: personal sites, company about pages, current team rosters. Nothing invented. Every affiliation current.
That is not the good news it looks like. The engines that fabricated in May were Gemini, Perplexity and Bing. ChatGPT and Claude fabricated nothing in the original run either. This time I reached ChatGPT, Claude and DeepSeek.
Two of the three engines I tested were already clean, and the three that invent people are the three I couldn't reach.
So hallucination is unmeasured in this run, not improved. The honest state of the fabrication question is exactly where May left it.
One softer failure did show up. Claude gave Lily Ray's title as "VP of SEO Strategy & Research"; her own site says "VP, SEO & AI Search at Amsive." Right person, right company, wrong title. Not an invented human, but the same family of error in a form that's much harder to catch, because nothing about it looks wrong.
What I'd actually do with this
The advice from the original study survives, with one addition and one demotion.
The newsletter finding stands, and it's the only move that puts you in both the recalled and the retrieved lists. That's still the whole game.
New: aim at the models that read carefully, not the ones that skim. Publishing a method or a number someone can replicate is what got names past Claude's filter. Being in a "top experts" roundup (or leaning on llms.txt and hoping) is what got names into DeepSeek's single-run output and then out of it again on the next attempt.
Demoted: chasing the roundups. Getting into an agency listicle looks like progress and produces a 1/9. It'll show up in someone's screenshot of a single ChatGPT session and vanish by the time anyone checks.
Method, so you can break it
Three engines, three runs each, fresh session per run, same prompt every time: "Name the 8 best AI SEO experts to follow right now, each with their company." Run August 4, 2026. Names and affiliations then verified individually against primary sources rather than against each other.
If you want to reproduce this on your own niche, the only part that matters is the repetition. One run per engine will give you a table that looks exactly as authoritative as this one and tells you much less. Ask three times. Count.
If you want to automate the spot-checking, the weekly monitoring stack and the AI SEO automation guide cover the machinery.
Limitations, plainly
Three engines, not five, and not the three that historically fabricate. Three runs is enough to expose noise and nowhere near enough to characterise it properly. All three engines were reached through command-line tools rather than the chat interfaces most people actually use, and the tooling itself turned out to be a variable: several runs stalled entirely until I disabled an agent wrapper that was quietly doing tool calls instead of answering. Single niche, single question, single day.
Next run: five engines if I can get the browser-based ones back, five runs each, model versions recorded at the point of asking. And I'll wait for all of the data before writing the conclusion, which I evidently need telling.
FAQ
How do I get cited by ChatGPT?
Have content that ChatGPT's web browsing tool can read (publish on a crawlable site), have it cited by third-party sources that ChatGPT also reads, and write under a consistent personal brand. There's no submission process — citations are emergent.
How do I get cited by Perplexity?
Same plus a recency dimension. Perplexity weights freshly-updated and well-cited recent content. Update your evergreen posts quarterly; that bumps the dateModified signal Perplexity uses.
Does paying for placements help with AI citations?
Indirectly and only slightly. If a paid placement lives on a high-authority site that AI engines read regularly, the placement might increase citation frequency for queries adjacent to that placement's topic. But the cost-to-impact ratio is poor compared to earning citations through actual work.
How often does AI citation change?
For a given query, citation patterns shift slowly — weeks to months. Major model updates (e.g., GPT-5 launch, Claude version bumps) can produce sudden shifts. For high-volume topics, expect noticeable change every 60-90 days.
Can I track which AI engines cite me?
Partially. Perplexity exposes citations in its UI explicitly. ChatGPT and Claude reveal sources when they cite (and don't always cite). Gemini and Bing Copilot are more opaque.
There's no equivalent of Google Search Console for AI engines yet. The closest tool is to manually spot-check key queries weekly.
What's the difference between AI citation and Google ranking?
Different audiences, different signals. Google rewards on-page optimization, backlinks, and click signals. AI engines reward entity clarity, cross-source consistency, citation graphs, and dateModified freshness. The overlap is real but not total — you can rank well on Google and not be cited by AI engines, and vice versa.
Why do AI models give different answers to the same question?
Mostly because of what they read at the moment you ask. In this study the engine that searched two listicles returned a different list on every attempt, while the engine answering purely from training data returned six of the same eight names three times running. Recall is frozen; retrieval changes with whatever ranks that day. Sampling randomness adds some wobble on top, but the retrieval effect is much larger.
How many times should you run an AI citation study?
At least three per engine, and report how often each name appeared rather than whether it appeared. Across nine runs here, twelve people were named exactly once and never again. A single run reports all twelve as findings and gives you no way to separate them from the names that came back eight times out of nine.
Does getting into a "top experts" listicle help you get cited by AI?
Briefly and unreliably. Listicle placements showed up in single runs of the engine that skims sources, then vanished on the next attempt. One model read an agency's roundup and recommended that agency's own founder, which is the mechanism working exactly as the people building those pages intend. It didn't survive a re-ask.
Which AI engine is most reliable for questions about people?
On this evidence, the ones that read widely and say what they're discounting. The engine that read eight sources and openly flagged which lists looked gamed produced the most stable answers across repeats. The one that read two sources produced a different answer every time. Neither invented anybody, but one was reliably repeatable and one was not.
Do AI engines make up people?
Some do. In the May run one engine invented an expert outright, complete with job title and agency, and three of five fabricated something. Nothing was fabricated across the nine August runs, but those covered engines that were already clean in May, so that's a coverage gap rather than an improvement. Verify names against a primary source before repeating them.
