Inside Our AI Site Audit: A Technical Specification for Modern SEO Scoring
Why we rebuilt our scoring engine
Most site audit scores are either too generic (a single “health” number) or too opaque (black-box algorithms you can’t explain to stakeholders). In 2026, with AI crawlers, answer engines, and schema expectations evolving rapidly, we needed a scoring system that is:
- Transparent: Every axis and deduction is documented and reproducible.
- Data-driven: Leverages industry-standard signals from DataForSEO’s On-Page API.
- Future-proof: Explicitly accounts for AI bot access and answer-ready content—not just traditional SEO.
Below is the full technical specification of our v2 AI Site Audit scoring logic.
The five scoring axes (100 points total)
We score sites across five weighted axes. Each axis is computed as a float, summed, then rounded once at the end. If the total exceeds 100, we log a warning instead of silently clamping.
| Axis | Max points | What it measures |
|---|---|---|
| Crawlability | 20 | Can search engines and AI agents reliably discover and index your pages? |
| AI Bot Access | 15 | Are major LLM crawlers (GPTBot, Claude-User, PerplexityBot, etc.) allowed via robots.txt? |
| Structured Data | 15 | Is valid schema markup present, and does it include high-value types? |
| Answer-Ready | 25 | Are pages structured to be extracted and cited by answer engines (title, H1, FAQ patterns, lists)? |
| Content Quality | 25 | Are pages free of critical issues (duplicates, errors, thin content)? |
Sample size note: If fewer than 5 pages are successfully fetched, we append: “Score based on N pages — treat as directional only.”
1. Crawlability (20 pts) — DataForSEO-backed
We use the DataForSEO On-Page API (/v3/on_page/pages/ and /v3/on_page/summary/) to fetch per-page signals for up to 50 priority pages (sorted by estimated traffic).
Signals extracted per page:
meta.title,meta.description,meta.htags.h1meta.canonical,meta.robots_info.is_crawlable,meta.robots_info.is_indexedmeta.content.plain_text_word_countchecks.is_4xx_code,checks.is_5xx_code,checks.duplicate_title,checks.duplicate_description,checks.duplicate_content,checks.low_content_rate
Scoring logic:
- Start at 20
- –6 if sitemap is missing/unreachable (
summary.sitemap_status) - –5 if robots.txt is missing/returns error
- –5 × (noindex_pages / total_pages) for proportional noindex penalty
- –4 × (non_crawlable / total_pages) for proportional non-crawlable penalty
- Floor at 0
This mirrors DataForSEO’s own onpage_score methodology, which weights issues by prevalence rather than binary pass/fail.
2. AI Bot Access (15 pts) — Custom logic
Traditional audits ignore AI crawlers. We don’t. We parse robots.txt and check allowances for:
AI_BOTS list (2026):GPTBot, ChatGPT-User, Claude-User, Claude-SearchBot, Google-Extended, PerplexityBot, Bingbot, Amazonbot, Applebot, Meta-ExternalAgent, Bytespider, CCBot
Key fixes:
- Corrected a bug where consecutive
User-agentgroups were misparsed, causing restrictive rules to appear permissive. - Unreachable robots.txt = 0 points (cannot verify = cannot pass).
- Docstring updated to explicitly name
Claude-UserandClaude-SearchBot.
3. Structured Data (15 pts) — Baseline + bonus model
We moved away from penalizing “basic” schema. Instead, we use a baseline + bonus model aligned with DataForSEO’s treatment of structured data as a hygiene factor.
Per-page weighting:
- 1.0 if
structured_data.typesincludes high-value types:FAQPage,Article,Product,HowTo,Review,VideoObject,NewsArticle,Recipe,Event - 0.6 if only basic types:
WebSite,BreadcrumbList,SiteNavigationElement,SearchAction - 0.0 if no valid schema (empty, malformed, or unparseable JSON-LD)
Axis score:round(15 × sum(page_weights) / total_pages)
Note:
structured_data.count > 0alone is insufficient—at least one recognizable@typemust be present.
4. Answer-Ready (25 pts) — Semantic signals, not just schema
Answer engines (Google AI Overviews, Perplexity, Bing Chat) extract content based on visible structure—not just JSON-LD. Our scoring reflects this.
Required signals (page fails if any missing):
titlepresentH1presentmeta descriptionpresent
Optional semantic signals (each adds to quality tier):
- FAQ signal (priority order):
(a)checks.has_faqfrom DataForSEO (if present);
(b) Visible Q&A pattern: question heading/bold text ending in?+ 40+ word answer paragraph;
(c)<details>/<summary>blocks.
→ Credit if any of the three detect FAQ content. Never use schema types alone. - List structure:
ul/olwith 3+ items - Content depth:
plain_text_word_count >= 200(skipped for utility pages:/contact,/privacy,/terms,/sitemap,/login,/cart,/404)
Page quality tiers:
- 3 required + 3 optional → 1.0
- 3 required + 2 optional → 0.85
- 3 required + 1 optional → 0.70
- 3 required + 0 optional → 0.55
- Any required missing → 0.0
JS-rendered pages: If HTML has <100 words and an empty root <div>, we flag as JS-rendered and exclude from scoring (not penalized).
Axis score:round(25 × sum(page_quality_weights) / total_pages)
5. Content Quality (25 pts) — Template-normalized thin content
We replaced the rigid “200-word” rule with DataForSEO’s checks.low_content_rate, which normalizes by page template—so utility pages aren’t unfairly penalized.
Problem page definition:
A page is problematic if any of: duplicate_title, duplicate_description, duplicate_content, low_content_rate, is_4xx_code, is_5xx_code. Fetch-failed pages count as problems.
Partial credit:
- 1 issue → loses 25% of contribution
- 2 issues → loses 75%
- 3+ issues → loses 100%
Axis score:round(25 × quality_score / total_pages)
Caching and performance
To balance freshness with cost/rate limits:
- Full DataForSEO API responses are cached in MongoDB under the client record with a timestamp.
- 24-hour TTL: We skip re-fetch if a cached result exists from within the last 24 hours.
force_refreshflag: Bypasses cache for on-demand re-audits.
This follows DataForSEO’s recommended pattern for scalable audits.
Test coverage (selected cases)
Our unit tests validate edge cases critical to scoring integrity:
- ✅ Page with only
WebSiteschema → 0.6 weight in Structured Data - ✅ Page with
FAQPageschema → 1.0 weight - ✅
/privacypage with low word count → not penalized in Content Quality - ✅ Fetch-failed page → counted as problem in Content Quality
- ✅ Unreachable
robots.txt→ 0 on AI Bot Access - ✅ Visible Q&A pattern (question heading + answer paragraph) → earns FAQ credit in Answer-Ready
- ✅ JS-rendered page → excluded from word count gate
- ✅
checks.has_faqabsent → handled gracefully without error
How this compares to industry standards
Our methodology aligns with leading SEO platforms (Screaming Frog, Sitebulb, Ahrefs) and DataForSEO’s own onpage_score calculation: docs.dataforseo
- Axis-based scoring with weighted dimensions
- Proportional deductions for noindex/non-crawlable pages
- Template-normalized thin content detection
- Single-round rounding (no per-axis clamping)
- Sample-size transparency for small crawls
Where we differ (and lead):
- Explicit AI bot access scoring (12 major LLM crawlers)
- Answer-ready semantic signals (visible Q&A, lists, depth) without over-relying on schema
- Baseline + bonus structured data model (no penalty for foundational schema)
Ready to see your score?
Run a free AI site audit on your domain and get a breakdown across all five axes—with actionable fixes prioritised by impact.
Start your audit for free


References & further reading
- DataForSEO On-Page API documentation
- How DataForSEO calculates
onpage_score - 120 OnPage API metrics explained
- Best practices for technical blog formatting docs.dataforseo