How Much Does Your Own Site Actually Matter for AI Citations?
Executive summary
AI answer engines now pick which brands get named when a buyer asks for options. Most marketing teams chase press and mentions to get picked. They assume their own website counts for very little.
The biggest study we have says the opposite. Yext looked at 6.8 million citations. 86% of them pointed at things the business already owns: its own website (44%) and its listings (42%). Even on questions that named no brand at all, a company’s own pages took nearly 60%.
That turns this into a budget question. For most firms, most AI visibility sits in space they already own. Fixing your own pages costs less than earning coverage does. And it keeps getting pushed down the list, because of studies that asked a much narrower set of questions than your buyers ask.
This article shows who that holds true for, who it doesn’t, and how to tell in about ten minutes.
How an AI answer gets made
1
A buyer asks
“Who should I use for this?”
2
It reads a few pages
A small handful, not the whole web
3
It names a few brands
And credits the pages it used
4
86%
of those pages are ones you control
Where the credited pages sit
6.8M citationsYour website 44% · your listings 42% · reviews and social 8% · news, forums and other 6%.
Which makes it two jobs, not a mystery
Your own pages
State your facts clearly, where a machine can read them
Your listings
Claim them, and keep every detail accurate
The other 14%
Reviews, forums, news. Influence, but no control
Yext, 6.8 million AI citations across ChatGPT, Gemini and Perplexity, 2025.
Key takeaways
- You control more of this than you’ve been told. Yext studied 6.8 million AI citations. 86% pointed at sources the brand manages itself: its own website (44%) and its listings (42%).
- Owned pages win the query that’s supposed to kill them. Questions with no brand name in them are meant to favor outside sources. On the factual ones, first-party sites and local pages still took nearly 60% of Yext’s citations.
- The studies that disagree measured a narrow slice. Their prompts were product rankings in three categories: consumer electronics, cars and software. The finding is real, but the window is small.
- Your query mix decides your budget. If customers reach you by asking about location, hours, coverage, eligibility or specifications, most of your outcome sits on pages you own. If they only reach you through “best tool for X” lists, it doesn’t.
- On-page work pays, within limits. The Princeton GEO study measured 30–40% more AI visibility from added structure, statistics and citations. Keyword stuffing scored below doing nothing at all.
You’ve probably read the opposite. The claim that AI search ignores your site in favor of earned media is true for one specific kind of question in a handful of consumer categories, and it’s been stretched into a rule that doesn’t fit most companies. The 86% figure holds across retail, financial services, healthcare and food service, which between them describe how most businesses actually get found.
This post is for marketers and executives deciding where the next unit of budget goes: your own pages, your listings, or a push for third-party coverage. It works through what each study actually measured, then gives a segment-by-segment answer on who should invest in the real estate they control. It doesn’t cover the technical mechanics of making pages readable to AI agents in the first place. We wrote a separate guide on that.
Two terms before the numbers. Generative engine optimization (GEO) is the work of getting cited by AI answer engines rather than ranked by search engines. Earned media is anything written about you on a page you don’t control: press coverage, analyst notes, review sites, comparison listicles, forum threads.
In this article
- The biggest dataset says you control most of what gets cited
- What the earned-media studies actually measured
- Which businesses should invest in the real estate they control
- Where off-site signals genuinely do the work
- Your Google rankings are not your AI citations
- What owned real estate actually buys you
- How to split the budget
- Where SiftServe fits
- Frequently asked questions
- Where to start
The biggest dataset says you control most of what gets cited
Yext pulled 6.8 million citations from 1.6 million queries per model across ChatGPT, Gemini and Perplexity, running from 1 July to 31 August 2025 and spanning 20,820 unique citation domains. The query set is the part to watch. They built it to cover four intent quadrants, branded and unbranded crossed with objective and subjective, across retail, financial services, healthcare and food service.
| Source type | Citations | Share | Who controls it |
|---|---|---|---|
| First-party websites | 2.9M | 44% | You, completely |
| Listings and profiles | 2.9M | 42% | You, through claimed and maintained entries |
| Reviews and social | 545K | 8% | You, indirectly |
| News, forums, other | — | 6% | Nobody, from your side |
The number to remember: across 6.8 million citations, 86% came from a brand’s own website and its listings. Only 6% came from sources no one on your side can reach.
Eighty-six percent sits in the top two rows. Those are pages you can edit this afternoon. Only 6% of citations landed somewhere you’ve no route into at all, and forums specifically, the Reddit effect that dominates GEO commentary, came in at around 2% once location and query intent were applied.
The finding that should reset the conversation is the quadrant split. Unbranded objective queries, the ones with no brand name in them, are supposed to be where your own site is least credible and earned media takes over. In Yext’s data, first-party websites and local pages supplied nearly 60% of citations on exactly those queries. When someone asks a question with a factual answer, the engine goes to the entity that holds the fact. That’s you.
Who gets cited depends on what was asked
Share of AI citations by source type, across three query sets.
All location-scoped queries
Yext · 6.8M citationsYour website 44% · your listings 42% · reviews and social 8% · news, forums and other 6%.
Unbranded, objective queries only
YextFirst-party sites and local pages take ~60% — on the queries that supposedly favour earned media.
Unbranded product-ranking prompts
arXiv · softwareBrand-owned falls to 26.7%; earned media takes 72.7%. This is the “best X for Y” shape.
Sources: Yext (rows 1–2), arXiv:2509.08919 (row 3). Row 1 is the study’s four control tiers; rows 2 and 3 are two-way splits.
Sources: Yext (rows 1–2), arXiv:2509.08919 (row 3). Row 1 percentages are the study’s four control tiers; rows 2 and 3 are two-way splits.
Industry by industry, the pattern holds, though the winning surface shifts:
| Industry | Largest single citation source | Share |
|---|---|---|
| Retail | First-party websites | 47.6% |
| Financial services | Brand-owned websites | 48.2% |
| Healthcare | Listings (WebMD, Vitals and similar) | 52.6% |
| Food service | Listings | 41.6% (reviews and social add 13.3%) |
Source: Yext AI citations research, Aug 2025.
Healthcare and food service are worth reading carefully. Your own website loses the top spot in both, but not to earned media. It loses to listings, which you also control. A claimed, accurate, complete profile is owned real estate that happens to sit on someone else’s domain.
The engines differ too, and the difference matters if you know where your traffic comes from. Gemini favors websites (52.1% of its citations). OpenAI leans on listings (48.7%). Perplexity spreads across a wider mix including MapQuest and TripAdvisor. If ChatGPT is your dominant AI referrer, and for most sites it is, your listings deserve as much attention as your homepage.
What the earned-media studies actually measured
Plenty of research points the other way, and it’s worth being precise about its scope rather than dismissing it.
Muck Rack’s December 2025 report found 82% of links cited by LLMs came from earned media, with journalistic sources holding just under 25% of all cited links. AirOps’ 2026 State of AI Search covered more than 21,000 brands and put 85% of brand mentions on third-party pages. Nearly 90% of those sat in listicles, comparisons or reviews. Both point the same way.
The most rigorous version is academic. Chen, Wang, Chen and Koudas ran 1,000 consumer ranking prompts across ten categories, then compared what four AI engines cited against what Google returned for the same query:
| Vertical (US) | Google: brand-owned | Google: earned | AI search: brand-owned | AI search: earned |
|---|---|---|---|---|
| Consumer electronics | 32.9% | 51.7% | 22.1% | 92.1% |
| Automotive | 39.5% | 45.1% | 18.1% | 81.9% |
| Software products | 43.7% | 45.4% | 26.7% | 72.7% |
Source: arXiv:2509.08919, Sept 2025. Figures shown for GPT; social content fell to roughly zero across all AI engines tested. Canadian queries put brand-owned at 22.1–30.9%.
That is a real effect and the numbers are sound. Now look at what produced them. One thousand prompts. Three verticals: consumer electronics, automotive, software. Every query a product ranking request of the “best laptop under $1,000” shape. Those are the three categories on earth with the deepest professional review ecosystems, asked in the one format that explicitly requests a comparison across brands. An engine answering “best laptop under $1,000” is being asked to do something no single manufacturer’s site can do, so it goes to the people who compare.
Yext ran 1.6 million queries per model across four industries and four intent types. The arXiv team ran 1,000 in one intent type across three verticals. Both results are valid within their scope. Only one of them describes the range of questions most businesses actually get asked.
Two more caveats belong here, in both directions. Muck Rack’s methodology isn’t published in full, and AirOps measures brand mentions rather than citations, which is a looser signal. On the other side, Yext sells listings management, and its query set is location-scoped, so it is sampling the conditions where owned real estate performs best. Nobody in this field is a disinterested party. What you can do is match the query set to your own.
Which businesses should invest in the real estate they control
Here is the practical version. Find the row that matches how customers actually find you.
| Your situation | Where citations come from | Owned-real-estate priority |
|---|---|---|
| You have physical locations (retail, restaurants, clinics, branches, dealerships) | Listings and local pages dominate: 41.6–52.6% in Yext’s food service and healthcare data | Highest. Claimed, complete, consistent listings plus location pages are close to the whole game. Earned media barely moves it. |
| You sell a service defined by eligibility, coverage or process (insurance, lending, legal, healthcare, logistics) | Brand-owned websites led financial services at 48.2% | Highest. Only you can state your terms, rates, coverage areas and requirements. Engines have nowhere else to go for them. |
| You have branded demand (people already search your name) | Branded queries lean on your site, your docs and your listings | High. Every branded question is yours to lose. An unreadable site means the engine answers from a stale third-party summary. |
| You sell something with specifications (equipment, components, B2B hardware, technical software) | Your spec pages, docs and pricing are the primary source; roundups quote them | High. Your documentation is the raw material for the comparisons others write. |
| You’re in a category with heavy review-media coverage (consumer electronics, cars, mainstream SaaS) | Earned media took 72.7–92.1% of AI citations in the arXiv tests | Moderate. Fix the floor, then spend on being covered and compared. |
| You’re new with no brand awareness, competing on “best X” shortlists | Listicles, comparisons, journalism and video | Lowest, for now. Nothing on your own domain gets you into a shortlist you’re absent from. Earn the mentions first. |
Most businesses are in the first four rows. Local, regional, professional-services, B2B and specification-driven companies make up the overwhelming majority of firms with a website, and for all of them the owned real estate is the majority of the answer. The earned-media-first advice that dominates GEO writing was derived from the bottom two rows and is being sold to everyone.
There’s also an asymmetry in effort worth naming. Getting cited in a comparison listicle means persuading someone else to write about you, on their schedule, with an outcome you don’t control. Making your own pages complete, accurate and machine-readable is work you can finish. When both routes lead to citations, the one you can finish should not be the one you postpone.
Test your own query mix in ten minutes
Don’t take a study’s word for which row you’re in. Write down the ten questions a customer would actually type before buying from you, then sort them:
- Does the question name you? Branded questions are yours to win or lose on your own pages and listings.
- Does it have one correct factual answer? Hours, coverage, eligibility, price, specification, compatibility. If yes, you’re the source an engine should reach for, branded or not.
- Does it ask for a ranked comparison across vendors? “Best X for Y”, “top tools for Z”. These are the earned-media queries.
Then run each one through ChatGPT, Gemini and Perplexity and record which domains get cited. If most of your ten fall into the first two buckets, you’re in the top four rows of the table and your budget belongs in owned real estate. If most fall into the third, you’re in the bottom two. The split is usually obvious after ten questions, and it’s worth redoing quarterly.
Where off-site signals genuinely do the work
None of this makes off-site signals optional. They decide something different: whether an engine considers you at all for a question that doesn’t name you.
Ahrefs measured this across 75,000 brands on ChatGPT, Google AI Mode and AI Overviews, filtering to domains rated 40+ with at least 800 monthly searches:
| Signal | Correlation with AI visibility | Where it lives |
|---|---|---|
| YouTube mentions | 0.737 | Off-site |
| YouTube mention impressions | 0.717 | Off-site |
| Branded web mentions | 0.66–0.71 | Off-site |
| Branded anchor text | 0.628 | Off-site |
| Domain rating | 0.266–0.326 | Mixed |
| Number of pages on your site | 0.194 | On-site |
Read that as a measure of brand-level discoverability, not page-level citation. It says a brand people talk about gets surfaced more often, which is unsurprising and not directly actionable in a quarter. The genuinely useful row is the last one. Publishing more pages correlated 0.194, which Ahrefs describes as almost no relationship. If your GEO plan is a programmatic content push, that number is the one to sit with.
SE Ranking’s model points the same way on volume. It’s built on 129,000 domains and 216,524 pages across 20 niches, and it turned up a useful negative result: llms.txt showed negligible impact on ChatGPT citation likelihood. Depth per page mattered instead. Articles over 2,900 words averaged 5.1 citations against 3.2 for those under 800.
Your Google rankings are not your AI citations
Whichever camp you land in, one assumption needs retiring: that winning SEO buys you GEO for free.
- Across 15,000 long-tail queries in July 2025, Ahrefs found only 12% overlap between AI assistant citations and Google’s top 10 results.
- For Google’s own AI Overviews the overlap is higher but falling fast: 38% of cited pages rank in the top 10, down from 76% a year earlier, across 863,000 keywords and 4 million AI Overview URLs.
- seoClarity’s analysis of ChatGPT’s 1,000 most-cited URLs found 25% of the top 100 have zero organic visibility in Google.
These are two different systems drawing on two different source pools. A page that ranks badly can still be the page an engine quotes, which is good news if your site is well-organized but young.
One caution before you build a plan on any single number in this post. Semrush tracked 230,000 prompts across 13 weeks and watched Reddit’s share of ChatGPT responses drop from roughly 60% to roughly 10%. That took one month. Source mixes move fast enough that any citation study really only describes the month it was run.
What owned real estate actually buys you
The Princeton GEO study (KDD 2024) is the cleanest evidence on how much on-page work is worth. Across GEO-bench, a benchmark of 10,000 queries, the researchers tested nine content edits and measured the change in citation rate:
| On-page edit | Relative change in visibility |
|---|---|
| Add quotations from credentialed sources | +41% |
| Add specific statistics | +33% |
| Fluency rewrite | +30% |
| Cite reliable external sources | +29% |
| Add technical terms | +19% |
| Add authoritative language | +18% |
| Easy-to-understand rewrite | +15% |
| Add unique words | +6% |
| Keyword stuffing | −8% |
Individual edits at the top clear 40%, but the durable band across domains, and the figure we quote, is 30–40% more AI visibility from added structure, statistics and citations. Note what failed. Keyword stuffing, the reflex of two decades of SEO, scored below doing nothing at all.
One honest limit on that number. Those gains were measured on content already in the retrieval set, the pool of documents an engine pulls before it writes an answer. The study changed how often a retrieved document got cited, not whether it entered the pool. For the business types in the top four rows of the table above, you’re already in the pool for the questions that matter, because you’re the entity the question is about. For the bottom two rows, entry is the problem and on-page edits won’t solve it.
There’s a second return on the same work. Journalists writing roundups, Reddit users answering questions and analysts building comparison tables all read your site before they write about you. Content that is hard for an AI agent to parse is usually hard for a human researcher to quote. Your owned pages feed the earned media that the other studies are measuring.
How to split the budget
- Make your own pages machine-readable first. Crawl access, content present in the raw HTML, structured data, answer-first sections, specific numbers with sources. This is a bounded project with an end date, and it’s the precondition for every other line item.
- Claim and complete every listing. Listings supplied 42% of Yext’s citations and led outright in healthcare and food service. For a business with locations this is the highest-return work available, and it’s usually the most neglected.
- Write the pages that answer factual questions. Coverage areas, eligibility, hours, pricing, specifications, comparisons against named alternatives. Unbranded objective queries drew nearly 60% of citations from first-party sites, and this is the content that wins them.
- Don’t scale pages for their own sake. Page count correlated 0.194 with AI visibility. Depth per page is the lever; volume of pages isn’t.
- Spend on earned media in proportion to your row. If you’re in a review-heavy category or have no brand awareness, this moves to the top of the list. If you’ve locations, branded demand or a specification-driven product, it’s a second-order investment.
- Track branded and unbranded prompts separately, and re-measure quarterly. Averaging them produces a number that describes none of your actual situations, and source mixes shift too fast for an annual review.
Where SiftServe fits
SiftServe is web infrastructure for AI agents. We build and serve an AI-readable version of your site: semantic HTML, FAQs and structured data, generated by an AI agent, approved by you, served to AI and bot traffic at the edge while human visitors see your site unchanged.
That’s step one on the list above, done properly. It matters most for the businesses in the top rows of that table, the ones whose own pages are the source an engine should be reaching for. An unreadable page forfeits citations that were yours by default. On our own homepage, the sifted version cut 382,057 bytes down to 33,191 while carrying more readable text than the original. About 97% of that original download was delivery scaffolding rather than anything a model could read, which is roughly what we find on most sites we measure.
Our working hypothesis, stated before the data and tested in every pilot, is that sifted pages earn around 20% more citations and AI referrals than their unsifted baseline. It stays a hypothesis until pilot data confirms or kills it.
Frequently asked questions
Which businesses get the most out of investing in their own website for AI citations?
Businesses whose customers ask factual questions only they can answer. That covers anyone with physical locations, any service defined by eligibility or coverage (insurance, lending, legal, healthcare), any product with specifications, and any brand that already has people searching its name. In Yext’s data, first-party websites led retail at 47.6% and financial services at 48.2% of citations, and listings led healthcare at 52.6%. The businesses that benefit least are new brands with no awareness competing purely on “best X” shortlists.
Is 86% of AI citations really from sources brands control?
That’s Yext’s finding across 6.8 million citations, and the 86% is first-party websites (44%) plus listings (42%). The qualifier matters. Those were location-scoped queries across retail, financial services, healthcare and food service. It’s the largest published citation dataset and the closest thing we have to ordinary commercial questions, but it isn’t measuring product-ranking prompts in review-heavy consumer categories, where the share drops sharply.
Do unbranded queries always go to third-party sources?
No, and this is the most over-generalized claim in GEO writing. Yext found first-party websites and local pages supplied nearly 60% of citations on unbranded objective queries. Unbranded subjective queries, and unbranded product rankings in consumer electronics, automotive and software, are where earned media takes over: 72.7–92.1% in the arXiv tests. The distinction is whether the question has a factual answer you own.
Should I invest in listings or in my own website first?
Both are owned real estate, so the split comes down to your category and your dominant engine. Listings supplied 42% of Yext’s citations overall and led healthcare (52.6%) and food service (41.6%). If you’ve physical locations, start there. OpenAI leans on listings for 48.7% of its citations while Gemini favors websites at 52.1%, so if ChatGPT is your main AI referrer, listings deserve the first pass.
Does ranking well on Google get me cited by AI engines?
Partly, and less than it used to. Ahrefs measured 12% overlap between AI assistant citations and Google’s top 10 across 15,000 long-tail queries, and found Google’s own AI Overviews now pull only 38% of citations from top-10 pages, down from 76% a year earlier. seoClarity found 25% of ChatGPT’s 100 most-cited URLs have no Google organic visibility at all.
Where to start
Find your row in the table above. If you’ve locations, eligibility rules, specifications or branded demand, the majority of your AI citations are sitting in real estate you already own, and the work is to make it readable and complete rather than to chase coverage.
Start by checking whether your content survives the trip to an AI agent at all. That takes a minute. Our no-rebuild guide walks the whole checklist, and the free AI visibility checker will tell you whether the basics are in place. Then claim and complete your listings, write the pages that answer the factual questions in your category, and hold earned media for the queries that genuinely need it.
If you’re in the bottom two rows and can’t get onto the shortlist at all, that’s a different problem with a different fix. We worked through it in why your competitors appear in ChatGPT and you don’t.
- AI citations
- AI search
- first-party content
- generative engine optimization
- GEO research