AIERP.cloud Site audit – Technical Guide

Inside Our AI Site Audit: A Technical Specification for Modern SEO Scoring


Why we rebuilt our scoring engine

Most site audit scores are either too generic (a single “health” number) or too opaque (black-box algorithms you can’t explain to stakeholders). In 2026, with AI crawlers, answer engines, and schema expectations evolving rapidly, we needed a scoring system that is:

  • Transparent: Every axis and deduction is documented and reproducible.
  • Data-driven: Leverages industry-standard signals from DataForSEO’s On-Page API.
  • Future-proof: Explicitly accounts for AI bot access and answer-ready content—not just traditional SEO.

Below is the full technical specification of our v2 AI Site Audit scoring logic.


The five scoring axes (100 points total)

We score sites across five weighted axes. Each axis is computed as a float, summed, then rounded once at the end. If the total exceeds 100, we log a warning instead of silently clamping.

AxisMax pointsWhat it measures
Crawlability20Can search engines and AI agents reliably discover and index your pages?
AI Bot Access15Are major LLM crawlers (GPTBot, Claude-User, PerplexityBot, etc.) allowed via robots.txt?
Structured Data15Is valid schema markup present, and does it include high-value types?
Answer-Ready25Are pages structured to be extracted and cited by answer engines (title, H1, FAQ patterns, lists)?
Content Quality25Are pages free of critical issues (duplicates, errors, thin content)?

Sample size note: If fewer than 5 pages are successfully fetched, we append: “Score based on N pages — treat as directional only.”


1. Crawlability (20 pts) — DataForSEO-backed

We use the DataForSEO On-Page API (/v3/on_page/pages/ and /v3/on_page/summary/) to fetch per-page signals for up to 50 priority pages (sorted by estimated traffic).

Signals extracted per page:

  • meta.title, meta.description, meta.htags.h1
  • meta.canonical, meta.robots_info.is_crawlable, meta.robots_info.is_indexed
  • meta.content.plain_text_word_count
  • checks.is_4xx_code, checks.is_5xx_code, checks.duplicate_title, checks.duplicate_description, checks.duplicate_content, checks.low_content_rate

Scoring logic:

  • Start at 20
  • –6 if sitemap is missing/unreachable (summary.sitemap_status)
  • –5 if robots.txt is missing/returns error
  • –5 × (noindex_pages / total_pages) for proportional noindex penalty
  • –4 × (non_crawlable / total_pages) for proportional non-crawlable penalty
  • Floor at 0

This mirrors DataForSEO’s own onpage_score methodology, which weights issues by prevalence rather than binary pass/fail.


2. AI Bot Access (15 pts) — Custom logic

Traditional audits ignore AI crawlers. We don’t. We parse robots.txt and check allowances for:

AI_BOTS list (2026):
GPTBot, ChatGPT-User, Claude-User, Claude-SearchBot, Google-Extended, PerplexityBot, Bingbot, Amazonbot, Applebot, Meta-ExternalAgent, Bytespider, CCBot

Key fixes:

  • Unreachable robots.txt = 0 points (cannot verify = cannot pass).
  • Docstring updated to explicitly name Claude-User and Claude-SearchBot.

3. Structured Data (15 pts) — Baseline + bonus model

We moved away from penalising the “basic” schema. Instead, we use a baseline + bonus model aligned with DataForSEO’s treatment of structured data as a hygiene factor.

Per-page weighting:

  • 1.0 if structured_data.types includes high-value types: FAQPage, Article, Product, HowTo, Review, VideoObject, NewsArticle, Recipe, Event
  • 0.6 if only basic types: WebSite, BreadcrumbList, SiteNavigationElement, SearchAction
  • 0.0 if no valid schema (empty, malformed, or unparseable JSON-LD)

Axis score:
round(15 × sum(page_weights) / total_pages)

Note: structured_data.count > 0 alone is insufficient—at least one recognizable @type must be present.


4. Answer-Ready (25 pts) — Semantic signals, not just schema

Answer engines (Google AI Overviews, Perplexity, Bing Chat) extract content based on visible structure—not just JSON-LD. Our scoring reflects this.

Required signals (page fails if any missing):

  • title present
  • H1 present
  • meta description present

Optional semantic signals (each adds to quality tier):

  1. FAQ signal (priority order):
    (a) checks.has_faq from DataForSEO (if present);
    (b) Visible Q&A pattern: question heading/bold text ending in ? + 40+ word answer paragraph;
    (c) <details>/<summary> blocks.
    → Credit if any of the three detect FAQ content. Never use schema types alone.
  2. List structure: ul/ol with 3+ items
  3. Content depth: plain_text_word_count >= 200 (skipped for utility pages: /contact, /privacy, /terms, /sitemap, /login, /cart, /404)

Page quality tiers:

  • 3 required + 3 optional → 1.0
  • 3 required + 2 optional → 0.85
  • 3 required + 1 optional → 0.70
  • 3 required + 0 optional → 0.55
  • Any required missing → 0.0

JS-rendered pages: If HTML has <100 words and an empty root <div>, we flag as JS-rendered and exclude from scoring (not penalized).

Axis score:
round(25 × sum(page_quality_weights) / total_pages)


5. Content Quality (25 pts) — Template-normalized thin content

We replaced the rigid “200-word” rule with DataForSEO’s checks.low_content_rate, which normalizes by page template—so utility pages aren’t unfairly penalized.

Problem page definition:
A page is problematic if any of: duplicate_title, duplicate_description, duplicate_content, low_content_rate, is_4xx_code, is_5xx_code. Fetch-failed pages count as problems.

Partial credit:

  • 1 issue → loses 25% of contribution
  • 2 issues → loses 75%
  • 3+ issues → loses 100%

Axis score:
round(25 × quality_score / total_pages)


Caching and performance

To balance freshness with cost/rate limits:

  • Full DataForSEO API responses are cached in MongoDB under the client record with a timestamp.
  • 24-hour TTL: We skip re-fetch if a cached result exists from within the last 24 hours.
  • force_refresh flag: Bypasses cache for on-demand re-audits.

This follows DataForSEO’s recommended pattern for scalable audits.


Test coverage (selected cases)

Our unit tests validate edge cases critical to scoring integrity:

  • ✅ Page with only WebSite schema → 0.6 weight in Structured Data
  • ✅ Page with FAQPage schema → 1.0 weight
  • ✅ /privacy page with low word count → not penalized in Content Quality
  • ✅ Fetch-failed page → counted as problem in Content Quality
  • ✅ Unreachable robots.txt → 0 on AI Bot Access
  • ✅ Visible Q&A pattern (question heading + answer paragraph) → earns FAQ credit in Answer-Ready
  • ✅ JS-rendered page → excluded from word count gate
  • ✅ checks.has_faq absent → handled gracefully without error

How this compares to industry standards

Our methodology aligns with leading SEO platforms (Screaming Frog, Sitebulb, Ahrefs) and DataForSEO’s own onpage_score calculation: docs.dataforseo

  • Axis-based scoring with weighted dimensions
  • Proportional deductions for noindex/non-crawlable pages
  • Template-normalized thin content detection
  • Single-round rounding (no per-axis clamping)
  • Sample-size transparency for small crawls

Where we differ (and lead):

  • Explicit AI bot access scoring (12 major LLM crawlers)
  • Answer-ready semantic signals (visible Q&A, lists, depth) without over-relying on schema
  • Baseline + bonus structured data model (no penalty for foundational schema)

Ready to see your score?

Run a free AI site audit on your domain and get a breakdown across all five axes—with actionable fixes prioritised by impact.

Start your audit for free

References & further reading

  • DataForSEO On-Page API documentation
  • How DataForSEO calculates onpage_score
  • 120 OnPage API metrics explained
  • Best practices for technical blog formatting docs.dataforseo