SiftServe scoring methodology – CORE-EEAT Content Benchmarks

· Rafi

Key takeaways

This page covers: What is CORE-EEAT; How to read the scores; Dimension and total scores; Contextual Clarity (agent experience, GEO); Organization (agent experience, GEO); Referenceability (agent experience, GEO); Exclusivity (agent experience, GEO); Experience (human experience, SEO).

What is CORE-EEAT

CORE-EEAT is the scoring standard behind every audit score SiftServe shows: 80 checks across eight dimensions, each dimension scored 0–100. The first four dimensions (Contextual Clarity, Organization, Referenceability, Exclusivity) measure how well AI agents can find, parse, and cite a page; together they form the GEO score. The last four (Experience, Expertise, Authority, Trust) are Google's E-E-A-T framing of credibility for human readers and search; together they form the SEO score.

The rubric builds on the open-source CORE-EEAT content benchmark by Aaron He Zhu, with the agent-experience checks grounded in the Princeton GEO study (Aggarwal et al., KDD 2024) and the human-experience checks grounded in Google's Search Quality Rater Guidelines.

This post is the reference for anyone reading a SiftServe audit report, and for anyone who wants to grade their own content by the same standard. It lists every check in all eight dimensions, the scoring formulas, and the per-content-type weights. It does not cover the capture stats (page weight, tokens, coverage); those have their own methodology post.

How to read the scores

Every page is checked twice using the same 80-item standard. The first check is for the original page, and the second is for the sifted copy that AI agents read. We compare the agent-experience dimensions, which focus on what AI crawlers and assistants can find and understand, between the original and sifted versions. The human-experience dimensions, which cover credibility signals for readers and search ranking, apply to the original page, since human visitors continue to see it while the sifted copy is shown to AI crawlers.

Our own homepage is the standing example: on the agent-experience view, the original page scores 65 and the published sifted copy (v10) scores 76. Both numbers, along with how they were produced, are in the capture stats post.

A few checks are vetoes rather than points. An undisclosed affiliate relationship, for example, fails the page outright regardless of its other 79 answers.

Dimension and total scores

Weighted scoring by content type

Different page types earn trust differently, so the total is also computed as a weighted sum: Weighted Score = Σ (dimension_score × weight)

Dim / Page typeProduct reviewHow-to guideComparisonLanding pageBlog postFAQ pageAlternativeBest-ofTestimonial
C10%20%10%20%25%25%10%10%10%
O10%20%20%10%10%25%15%25%5%
R15%10%25%5%10%15%25%20%15%
E20%5%10%5%20%5%5%15%10%
Exp20%5%5%5%10%5%15%5%30%
Ept5%20%15%5%10%10%5%10%5%
A5%5%5%25%5%5%5%5%5%
T15%15%10%25%10%10%20%10%20%

Read down a column and the logic shows: a landing page lives or dies on Authority and Trust (25% each), a blog post on Contextual Clarity (25%), a testimonial page on first-hand Experience (30%).

Contextual Clarity (agent experience, GEO)

CheckExperienceWhat good looks like
Intent AlignmentAgentTitle promise = content delivery
Direct AnswerAgentCore answer in first 150 words
Query CoverageAgentCovers ≥3 query variants (synonyms, long-tail)
Definition FirstAgentKey terms defined on first use
Topic ScopeAgentExplicitly states what is and isn't covered
Audience TargetingAgentStates "this article is for…"
Semantic CoherenceAgentLogical flow between paragraphs, no jumps
Use Case MappingAgentDecision framework: when to choose A vs B
FAQ CoverageAgentStructured FAQ covering long-tail follow-ups
Semantic ClosureAgentConclusion answers the opening question + next steps

References: GEO: Generative Engine Optimization (Aggarwal et al., KDD 2024) · Google: creating helpful, reliable, people-first content

Organization (agent experience, GEO)

CheckExperienceWhat good looks like
Heading HierarchyAgentH1→H2→H3, no level skipping
Summary BoxAgentHas TL;DR or Key Takeaways section
Data TablesAgentComparisons and specs presented in tables
List FormattingAgentParallel items use bullet or numbered lists
Schema MarkupAgentAppropriate JSON-LD (Article/FAQ/HowTo/etc.)
Section ChunkingAgentEach section has single topic; paragraphs 3–5 sentences
Visual HierarchyHumanKey concepts bolded or highlighted
Anchor NavigationAgentTable of contents with jump links
Information DensityAgentNo filler; consistent terminology throughout
Multimedia StructureHumanImages/videos have captions and carry information

References: The RefinedWeb dataset: filtering web data for LLM training (NeurIPS 2023) · GEO: Generative Engine Optimization (Aggarwal et al., KDD 2024)

Referenceability (agent experience, GEO)

CheckExperienceWhat good looks like
Data PrecisionAgent≥5 precise numbers with units (%, $, ms)
Citation DensityAgent≥1 external citation per 500 words
Source HierarchyAgentPrimary sources first; ≥3 Tier 1–2 sources
Evidence-Claim MappingAgentEvery claim backed by evidence immediately after
Methodology TransparencyAgentSample size, steps, and criteria documented
Timestamp & VersioningAgentLast updated <1 year; version changes noted
Entity PrecisionAgentFull names for people/orgs/products; no "a company"
Internal Link GraphHumanDescriptive anchor texts forming topic clusters
HTML SemanticsAgentUses <article>, <figure>, <time>, <cite>
Content ConsistencyAgentData self-consistent; no broken links (404)

References: GEO: Generative Engine Optimization (Aggarwal et al., KDD 2024) · Google: how structured data works

Exclusivity (agent experience, GEO)

CheckExperienceWhat good looks like
Original DataAgentFirst-party surveys, experiments, or statistics
Novel FrameworkAgentNamed, citable original framework or model
Primary ResearchAgentOriginal experiments/surveys with documented process
Contrarian ViewAgentChallenges consensus with evidence
Proprietary VisualsHuman≥2 original infographics, charts, or diagrams
Gap FillingAgentCovers questions competitors don't
Practical ToolsHumanDownloadable templates, checklists, or calculators
Depth AdvantageAgentDeeper than competing content on same topic
Synthesis ValueAgentCross-domain knowledge combination (A+B=C)
Forward InsightsAgentData-backed predictions and trend analysis

References: Google: creating helpful, reliable, people-first content · GEO: Generative Engine Optimization (Aggarwal et al., KDD 2024)

Experience (human experience, SEO)

CheckExperienceWhat good looks like
First-Person NarrativeHumanContains "I tested" or "We found" + action verbs
Sensory DetailsHuman≥10 sensory words (smooth, heavy, bright)
Process DocumentationAgentStep-by-step process with timeline
Tangible ProofHuman≥2 original photos/screenshots with timestamps
Usage DurationHumanStates "after X months of use…"
Problems EncounteredAgentShares ≥2 real problems + solutions
Before/After ComparisonHumanShows change, improvement, or difference
Quantified MetricsAgentMeasurable experience data (time, cost, success rate)
Repeated TestingHumanMultiple tests or long-term tracking
Limitations AcknowledgedAgentStates "we only tested X scenario"

References: Google: E-E-A-T in the Search Quality Rater Guidelines · Google Search Quality Rater Guidelines (PDF)

Expertise (human experience, SEO)

CheckExperienceWhat good looks like
Author IdentityHumanByline + avatar + bio (>30 words)
Credentials DisplayHumanRelevant degrees, certs, years of experience
Professional VocabularyAgentAccurate industry jargon, no misuse
Technical DepthAgentParameters, thresholds, examples are actionable
Methodology RigorAgentAnalysis method is reproducible
Edge Case AwarenessAgentDiscusses ≥2 exceptions or "when this doesn't apply"
Historical ContextHumanShows knowledge of the field's evolution
Reasoning TransparencyAgent"We chose A over B because…" with tradeoffs
Cross-domain IntegrationAgentConnects knowledge across fields
Editorial ProcessHuman"Reviewed by" or "Fact-checked by" labels

References: Google Search Quality Rater Guidelines (PDF)

Authority (human experience, SEO)

CheckExperienceWhat good looks like
Backlink ProfileHumanCited by authoritative sites (.edu, .gov, leaders)
Media MentionsHuman"Featured in" with media logos
Industry AwardsHumanDisplays relevant industry awards or recognition
Publishing RecordHumanConference talks, publications, patents
Brand RecognitionAgentBrand has search volume
Social ProofHumanAuthentic user testimonials with real details
Knowledge Graph PresenceAgentHas Wikipedia entry or Google Knowledge Panel
Entity ConsistencyAgentBrand/author info consistent across the web
Partnership SignalsHumanShows partnerships with authoritative organizations
Community StandingHumanActive and influential in professional communities

References: Google: E-E-A-T in the Search Quality Rater Guidelines · Google Search Quality Rater Guidelines (PDF)

Trust (human experience, SEO)

CheckExperienceWhat good looks like
Legal ComplianceHumanPrivacy Policy + Terms of Service present
Contact TransparencyHumanPhysical address or ≥2 contact methods
Security StandardsHumanSite-wide HTTPS, no security warnings
Disclosure StatementsAgentAffiliate links disclosed (veto if missing)
Editorial PolicyHumanContent standards and review process published
Correction & Update PolicyAgentHas corrections page or changelog
Ad ExperienceHumanAds <30% of page; no intrusive popups
Risk DisclaimersAgentYMYL topics have necessary disclaimers
Review AuthenticityAgentReviews show authenticity signals
Customer SupportHumanClear return policy, complaint channels, response SLA

References: Google Search Quality Rater Guidelines (PDF)

Further reading: how AI systems read your pages

Primary sources on how AI crawlers fetch and select content, and how generative engines choose what to cite:

Where the score fits

CORE-EEAT is one half of what a SiftServe audit reports; the other half is the capture stats (page weight, tokens, coverage), measured by their own published pipeline. Together they answer two questions: how much of a page an agent can read, and how well the page scores once read. If you want both run on your own site, before and after sifting, become a research partner.


Key facts

Visual content

Main image: A minimalist cream-colored landing page for “SiftServe.” with a small black-and-red logo at top left. Large bold text reads “Your next visitor isn’t human.” above “Make your site readable to AI agents.”, with red text at the bottom saying “ask → read → convert” and decorative horizontal black lines and colored dots on the right.

No significant hero imagery is shown; the page opens with a text-based blog article title.

About this company

SiftServe restructures existing website content into semantic HTML with explicit facts, FAQs, and Schema.org structured data. Customers review and approve each version before an edge worker serves it to AI and bot traffic, while human visitors and search-engine crawlers continue receiving the original site. SiftServe is a product of UAE-based NextON Consulting FZE.

Target customers:

Experience & expertise

People behind the company:

Products

Trust & authority

Backed by leaders from

Integrations

Contact

Social

Legal

Company

Pages

Frequently asked questions

What is SiftServe?
SiftServe is an infrastructure layer that creates an AI-readable version of an existing website and serves it to AI and bot traffic while leaving the human-facing site unchanged.
How does SiftServe work?
SiftServe crawls the customer's sitemap, compiles an editable company profile and voice guidelines, translates each page into semantic HTML with explicit facts, FAQs, and structured data, obtains human approval, and then serves the approved version to bots through an edge worker.
Does SiftServe change the website seen by human visitors?
No. Human visitors and search-engine crawlers such as Googlebot and Bingbot continue to receive the original website, while identified AI and bot traffic receives the approved sifted version at the same URL.
Does SiftServe generate new claims?
No. Its agent may restructure and rephrase what the source pages already assert, but every claim must trace to the original content and nothing publishes without customer approval.
How is SiftServe deployed?
An edge worker is deployed in front of the customer's domain, either in the customer's own Cloudflare account where available or on a supported CDN, without changes to the origin.
How much does SiftServe cost?
The service is currently offered as a free eight-week pilot on one domain, followed by a week-eight report that the customer uses to decide whether to continue.
How does SiftServe handle visitor and customer data?
Its edge software records no visitor IP addresses and sets no cookies, dashboards show visitor data only in aggregate, and model providers are accessed through business APIs whose terms do not permit training on customer content.
What is required to start a SiftServe pilot?
The stated prerequisites are a live domain, someone with authority to approve content, and a willingness to measure results.
How does SiftServe measure results?
It captures all bot requests and created deep analytical reports on their behaviour, which is available on your app. Every page receives a before-and-after audit score and per-page capture statistics, and the eight-week pilot tests citation and AI-referral performance against the unsifted baseline.
What does CORE-EEAT stand for?
CORE stands for Contextual Clarity, Organization, Referenceability, and Exclusivity — the four GEO dimensions that measure how well AI agents can find, parse, and cite a page. EEAT stands for Experience, Expertise, Authority, and Trust — Google’s E-E-A-T framing of credibility for human readers and search. Together they form an 80-item rubric scored across eight dimensions, each scored 0–100.
What counts as a good CORE-EEAT score?
Scores are most useful as a pair. SiftServe’s own homepage currently sits at 65 → 76 (original → sifted, version 10); an earlier, rougher version of both the page and the sift scored 33 → 44. The direction and size of the change matter more than any single absolute number, because many agent-experience checks assume structure that ordinary pages don’t carry.
Why do the GEO dimensions matter if my SEO is already strong?
Because they test different readers. A page can rank well for humans and still show an AI crawler a nearly empty document; the C, O, R, and E checks measure what survives a single no-JavaScript fetch. The capture stats post (https://siftserve.com/blog/how-we-measure-the-capture-stats) shows how large that gap gets in bytes and tokens.
Why are the CORE-EEAT weights different per content type?
Because readers trust different page types for different reasons. A testimonial page earns its keep through first-hand Experience (30%), a landing page through Authority and Trust (25% each). Grading every page on one flat rubric would reward blog-post structure everywhere, including where it doesn’t belong.
Can I score my own pages against the CORE-EEAT rubric?
Yes. Every check is listed in this post, and the underlying benchmark is open source (https://github.com/aaron-he-zhu/core-eeat-content-benchmark). Working through the 80 items by hand takes an hour or two per page; a SiftServe pilot runs the same rubric automatically on every page, original and sifted.
Does a sifted page always score higher than the original on CORE-EEAT?
On the agent-experience dimensions (C, O, R, E), that is the point of sifting, and the audit runs on every revision to verify it. Human-experience dimensions like Authority depend on off-page signals (backlinks, brand recognition) that sifting doesn’t change.
What are the CORE-EEAT scoring formulas?
GEO Score = (C + O + R + E) / 4; SEO Score = (Exp + Ept + A + T) / 4; Total Score = (GEO Score + SEO Score) / 2. A weighted score is also computed as Σ (dimension_score × weight), where weights vary by content type — for example, Contextual Clarity is weighted 25% for blog posts, while Authority and Trust each carry 25% for landing pages.

Sources