—hi-ai← Back to audit
— methodology

How we score your site

AI agents don't view the rendered UI the way a human does. They get a dump of the raw HTML inspector and reason about what your site is from there. The 34 checks in this audit measure the signals that dump actually carries — semantic HTML, structured data, accessibility attributes, and the files crawlers explicitly look for.

— two dimensions

Access, then answerability

The report splits into two scores because they answer two different questions. AI Access measures whether crawlers can find and read your site at all — robots.txt rules, a server-rendered HTML body, a sitemap, llms.txt, freshness for re-crawl. If a bot can't get in, nothing else matters. Sections 01 and 03 roll up here, plus the bot-allow signals in 05.

AEO Readiness — answer engine optimization — measures whether, once an engine is in, it can lift a clean, attributable answer and cite you. Entity identity (02), quotable content shape (04), the schema and freshness signals in 05 and 06, and the new Answerability section (07) all roll up here. Each check belongs to exactly one dimension, and the two maxes sum to the full point pool, so the headline number is just their combined average.

The three things that decide everything

  • HTML semantics. An <article> wrapped around a piece of content tells the model "this is the thing." A wall of <div>s tells it nothing. The same words score differently depending on the tags they sit inside.
  • Accessibility attributes. alt on images, aria-label on interactive elements, semantic heading order — every signal that makes your site usable for a screen reader makes it legible to an LLM. Accessible sites are AI-readable by construction.
  • Structured data. JSON-LD schema and OpenGraph meta tags act as context anchors. When the model has to decide whether you're a payments company or a software vendor, those tags are the evidence it cites.
— per-section rubric

What each section checks

01

AI Accessibility

AI crawlers check robots.txt before they fetch anything, then need a usable HTML body. If either step fails, no later check matters — the content is invisible to AI.

Signals measured
  • robots.txt allow/disallow rules for the major AI user agents (GPTBot, ClaudeBot, PerplexityBot, Googlebot, Google-Extended, Applebot-Extended, ChatGPT-User).
  • HTTP status of the homepage; presence of substantial text in the raw HTML body (not just after JavaScript runs).
  • Meta directives (`noindex`, `nofollow`, `noai`, `noimageai`) in <meta> tags or `x-robots-tag` response headers.
Why it's scored this way

AI agents do not render the page in a real browser the way a human visitor does. They fetch the raw HTML, parse it server-side, and reason about its structure. A single-page-app shell that renders content only after JavaScript executes registers as an empty page to most AI tools.

02

Brand Identity

AI models look for a machine-readable record that names who you are, what you do, and why you are credible. The signals are mostly invisible to humans — they live in JSON-LD, OpenGraph, and the document <title>.

Signals measured
  • JSON-LD `Organization` / `WebSite` / `Person` schema with `name`, `sameAs`, and identifying URLs.
  • OpenGraph + Twitter card metadata: `og:title`, `og:description`, `og:image`, `og:site_name`.
  • Document title and meta description shape (not too short, not over-stuffed).
  • Canonical URL, declared author, declared publisher.
Why it's scored this way

When you ask an AI "who is Stripe?", the model is largely citing the structured data that Stripe publishes about itself. Sites without this layer get cited less, get cited inaccurately, or get conflated with unrelated entities.

03

Content Readability

A handful of conventional files tell AI agents what is on your site and how to navigate it. They are cheap to publish and load-bearing for inclusion.

Signals measured
  • `/sitemap.xml` referenced from robots.txt and covering primary pages with valid `<lastmod>` dates.
  • `/llms.txt` (and optional `/llms-full.txt`) declaring a curated index of pages for AI consumption.
  • `<lastmod>` recency across the sitemap as a freshness signal.
Why it's scored this way

AI tools that respect these files use them to decide what to crawl and how to prioritize re-crawls. Missing or stale files mean the AI is guessing at your structure from links alone, which is slower and less reliable.

04

Quotability

AI models prefer content shaped like an answer: a short self-contained paragraph after a heading, with lists and tables breaking up dense prose. The same writing arranged poorly gets skipped.

Signals measured
  • Semantic HTML — `<article>`, `<section>`, `<main>`, `<h1>` through `<h4>`, not a wall of `<div>`s.
  • Paragraph length and sentence shape (short, declarative, scoped to one idea).
  • Presence of structured lists, tables, and inline emphasis on key terms.
  • A direct-answer first paragraph that contains the page's thesis, not a marketing intro.
Why it's scored this way

When the model decides what to quote, it picks the smallest self-contained chunk that answers the user question. A page where the answer is buried in paragraph six rarely makes it into a citation.

05

Platform Fit

Each AI platform weights different signals. A page can score well overall and still be invisible to one of them because of a single missing piece.

Signals measured
  • ChatGPT: needs GPTBot + ChatGPT-User allowed; favors JSON-LD and llms.txt.
  • Perplexity: needs PerplexityBot allowed; favors fresh `<lastmod>` and clear citations.
  • Claude: needs ClaudeBot allowed; favors structured content and explicit metadata.
  • Google AI Overviews: needs Googlebot AND Google-Extended allowed; favors `Article` / `FAQ` schema and OG images.
Why it's scored this way

A site can have great content and still be locked out of a single surface because of one robots.txt line or one missing schema block. This section scores each surface separately so the gap is named.

06

Site Hygiene

Two final checks: that brand identity reads the same way across every signal, and that your server can hand back a clean markdown version of any page when an AI asks for one.

Signals measured
  • Brand name consistency across `<title>`, JSON-LD `Organization.name`, OpenGraph `og:site_name`, and Twitter cards.
  • Content negotiation: does the page return useful content when `Accept: text/markdown` is requested?
Why it's scored this way

Inconsistent brand names confuse the model about which entity it is citing. A markdown response variant is a low-cost upgrade that AI tools increasingly prefer because it strips the rendering noise.

07

Answerability

Reading your page and wanting to quote it are different things. Once a crawler is in, these checks score whether your answers are shaped — and attributed — for an engine to lift them cleanly into an AI answer.

Signals measured
  • A `FAQPage` JSON-LD block with three or more question/answer pairs, scored separately from the homepage FAQ structure in Section 04.
  • Headings phrased as the questions users actually type — ending in "?" or opening with a question word.
  • Answer-first sections: the first paragraph under a heading answers it directly, in 80 words or fewer, without a windup.
  • Person-level attribution: a visible byline next to a date, plus a `Person` JSON-LD node with `sameAs` links.
  • Citable markers — percentages, dated figures, dollar amounts, named sources — over vague adjectives.
  • On-page freshness: a visible "Last updated" date or `<time datetime>` element (sitemap `lastmod` is scored in Section 03).
  • A pre-summarized "TL;DR" / "Key takeaways" block an engine can quote wholesale.
Why it's scored this way

An engine answering a user question wants the smallest attributable chunk that resolves the prompt. A page that frames its headings as questions, answers them immediately, backs claims with concrete figures, and names its author gives the engine exactly that — and gives it a reason to cite you by name rather than paraphrasing anonymously.

— sources

Where this rubric came from

The signals here are drawn from the public crawler documentation each AI vendor publishes (OpenAI's GPTBot and ChatGPT-User specs, Anthropic's ClaudeBot, Perplexity's PerplexityBot, Google's Googlebot and Google-Extended docs, Apple's Applebot-Extended), the schema.org vocabulary, the IETF robots.txt RFC 9309, the Web Content Accessibility Guidelines, and the emerging llms.txt convention from llmstxt.org. We weight each check by how much it materially shifts whether the page lands in an AI answer — not by how easy it is to test for. The full check list and weights live in the open-source code at apps/web/lib/hi-ai/.

hi-ai · audits · © 2026 KYA-OS
kya.vouched.id · kya@vouched.id