The five big AI engines barely agree on anything, comparison sites and Reddit run the show, and the “answer” changes almost every time you ask. Here’s what that means for your brand’s AI visibility, and what to do about it.
Everyone’s asking the same question right now: are we showing up in AI? So teams screenshot a ChatGPT answer, see their brand (or a competitor’s), and put it in a deck like it’s a ranking.
We wanted to know whether that screenshot means anything. So we ran the numbers.
Across two weeks, we put 80 Australian buying topics to the five services people actually use, Google, ChatGPT, Gemini, Claude and Perplexity, asked each one three different ways, and re-checked at the baseline, one day later and one week later. That’s 3,600 completed searches and 27,924 individual citations behind the answers.
Then we looked at which websites AI actually cites, how much it changes, and how differently each engine behaves. Some of it confirmed what we suspected. Some of it genuinely surprised us.
6 things we learned
The five engines barely agree. Ask the same question across all five, and only 8% of the cited sources are shared.
- Comparison sites and Reddit run the show. The four most-cited websites in Australian AI answers are Canstar, Reddit, Finder, and YouTube, not brands.
- The answer won’t sit still. Nearly half of an answer’s sources are gone the next time you look.
- AI isn’t a re-skin of Google. Even the closest engine shares only about 1 in 10 of its sources with the Google results page.
- Every engine has a personality. Perplexity cites 3.5× more sources than ChatGPT and leans on Reddit; Gemini writes in American spelling; only ChatGPT makes its citations show up in your analytics.
- One word changes the whole answer. Ask for the “best” and the comparison sites double while the hard numbers drop.
Let’s go through them.
1. The five AI engines barely agree
We measured overlap the simple way: for the same topic, how many of the cited websites two answers had in common. Same question, five engines, same moment, and only 8% of the cited sources were shared across the platforms.
It gets starker at the domain level. Of the 4,625 different websites cited across the whole study, 71% appeared on just one engine. Only 65 of them (1.4%) were cited by all five.

So “we show up in AI” isn’t one fact. It’s five different facts, and they mostly don’t line up.
What this means for you: treat each engine as its own channel. Being cited in ChatGPT tells you almost nothing about Gemini or Perplexity. If AI visibility matters to your category, you have to measure it on each platform your audience actually uses, not sample one and assume the rest.
2. Comparison sites and Reddit run the show
Here’s the one that should reframe how you think about “AI SEO.” The most-cited websites across all 3,600 Australian answers weren’t brands. They were intermediaries:
- canstar.com.au – cited in 14% of all answers
- reddit.com – 14%
- finder.com.au – 12%
- youtube.com – 11%
And it’s concentrated by category in a way that matters. In our study, Canstar appeared in 72% of finance answers and Finder in 68% of insurance answers. In finance and insurance, a comparison site is cited roughly 30% of the time, while in health and travel it’s around 1%.

Consumer categories tell a different story again: Reddit was the single most-cited source for both automotive and retail. Not a brand. Not a review site. A forum.
What this means for you: in money categories, your customer often meets a comparison site before they meet you, so earned presence on the sites AI already trusts can matter more than your own pages. In consumer categories, the conversation is happening in UGC. And in regulated categories like health, official sources dominate. There is no single playbook; there’s a playbook per category.
3. The answer won’t sit still
This is why a single screenshot is such a shaky thing to plan around.
We checked each question at three points in time. Across those checkpoints, 47% of the sources in an answer showed up only once. Only about a quarter held all the way through. Come back a week later and the answer has quietly swapped out roughly 40% of its sources, and usually added a few new ones on top.

So when you screenshot an AI answer, about half of what you’re looking at is a one-time appearance that may never happen again.
What this means for you: stop optimising around single answers. The ~25% of sources that do recur are the real signal, those are the pages and publishers worth investing in. The rest is noise. Measure the pattern that repeats, over time, not the moment you happened to catch.
4. AI isn’t a re-skin of Google
A comforting assumption floating around is: “we rank well on Google, so we’ll be fine in AI.” We tested it directly, comparing each engine’s cited sources against Google’s own results page.
Even the closest engine, Perplexity, shared only about 11% of its sources with Google. Claude shared just 6%.

The AI source universe is largely its own thing. Winning the search results page does not automatically put you in the AI answer.
What this means for you: audit AI visibility separately from search rankings. They draw on different sources, and a strong Google position is not a proxy for AI presence.
5. Every engine has a personality
We didn’t just count citations, we read 2,880 of the actual answers. The engines don’t just cite different sources; they behave like different writers.
- ChatGPT is the concise adviser: shortest answers (~270 words), easiest to read, talks to you, and asks for your postcode before recommending.
- Gemini is the formal encyclopedia: long, dense, the hardest to read, and it writes in American spelling (only 17% of its spellings were Australian, versus ~64% for the others) even though every question said “in Australia.”
- Claude is the structured guide: the most headings, the most emoji, opinionated.
- Perplexity is the dense analyst: the longest answers, cites 3.5x more sources than ChatGPT, and leans hard on Reddit (37% of its answers) and YouTube (31%).

And one finding with real practical application: only ChatGPT makes its citations measurable. In 90% of its answers, ChatGPT linked sources with a “utm_source=chatgpt.com” tag, so if it cites your page and someone clicks, it shows up in your analytics as ChatGPT. Gemini and Perplexity carry no in-body links at all, so their referrals are effectively invisible in standard reporting.
What this means for you: match the platform to the goal. Want measurable referral traffic today? ChatGPT is the only one you can actually count. Want to be part of the conversation Perplexity leans on? That’s Reddit and YouTube. Want to be described in Australian English? Anyone but Gemini.
6. One word changes the whole answer
We asked the same intent two ways: the plain keyword (“car insurance”) and the superlative (“best car insurance”). Adding one word reshaped the answer.
Asking for the “best”:
- roughly doubled the comparison sites cited (7% → 14% of sources)
- tripled the superlative language in the writing (“best,” “top,” “leading”)
- and quoted fewer actual numbers

Your customers don’t all ask the same way. “Car insurance” and “best car insurance” are effectively two different questions that surface two different source worlds, and the commercial-intent “best” version is exactly where the comparison sites win.
What this means for you: measure across the phrasings your customers actually use, not one preferred prompt. And know that “best”-style queries are a ranked-recommendation game, the content you need to be part of there is a shortlist, not a spec sheet.
The bigger picture
Put it together, and the headline is simple: AI visibility isn’t a ranking. It’s a moving system.
It changes with time, with wording, and with platform. It’s intermediated by comparison sites and forums more than by brands. And it barely overlaps with the Google results you already track.
That doesn’t make it unmeasurable, it makes the screenshot the wrong unit of measurement. A better AI-visibility review answers five questions instead:
- How often does your brand appear: across a real set of customer questions, not cherry-picked ones?
- Under which phrasings: including the “best”-style queries where intent lives?
- On which platforms: measured separately, because they don’t agree?
- Alongside which sources: the comparison sites, forums and publishers AI actually leans on?
- Does it persist: over repeated checks, or was it a one-time flicker?
Answer those and you stop chasing screenshots and start seeing the pattern that’s actually worth acting on.
How we did this
We built the study from 80 Australian buying topics across 10 industries (finance, insurance, home services, health, travel, automotive, real estate, legal, retail and B2B software). Each topic was asked three ways:
- a plain keyword,
- the same with “Australia,”,
- a “best…” comparison version
Run on Google, ChatGPT, Gemini, Claude and Perplexity, at a baseline, one day later and one week later.
That’s 3,600 completed searches and 27,924 normalised citations. We measured overlap as the share of cited websites two answers had in common.
What this doesn’t measure: accuracy, authority, sentiment or commercial impact. Low overlap doesn’t mean an answer is wrong, it means the evidence behind it changed. Category and phrasing differences are observed associations, not proven cause and effect. Named sources are reported by how often they were cited, not ranked by quality.
Where to from here?
If you want to know your own picture, who intermediates your category on each engine, how stable your visibility is, and where your content and earned-authority work should focus, that’s exactly what an AI Search Visibility & Citation Audit is for. It establishes a defensible baseline, then shows where to act.
Measure the pattern before you act on the screenshot.