Can AI read your website?
This scan answers that in about ten seconds: it checks the nine technical foundations where AI readability most often fails, from AI crawlers locked out in robots.txt (the most common silent mistake) through llms.txt and sitemap to Schema.org markup and meta tags. Every finding comes with the reason it matters and the way to fix it.
Only publicly accessible files are checked, nothing is stored. Whether AIs actually know and recommend your business is measured by the AI Visibility Check; the two tools complement each other: the outcome there, the technical cause here.
The scan
Enter a URL, done.
Takes about 10 seconds · no sign-up · nothing is stored
Result for
Extra credit · agent readiness · not part of the score
The head-start checks.
What is checked, and why it matters
Nine checks, weighted by impact. The most important one comes first: a website that locks out AI crawlers simply does not exist for AI assistants, no matter how good the rest is. Plus four extra-credit checks for agent readiness: they count as a bonus but never lower the score, almost nobody has these standards today, whoever does is ahead of the curve.
| Check | Weight | Why it matters |
|---|---|---|
| AI crawlers in robots.txt | 20 | Blocking GPTBot (ChatGPT), ClaudeBot (Claude) or PerplexityBot means not appearing in their answers. Many templates and firewalls block across the board, often without the owner knowing. |
| Schema.org markup | 15 | Structured data tells machines explicitly what the page is: a company, a service, an FAQ. A basic prerequisite for being cited correctly. |
| Title & meta description | 15 | The two lines every search engine and every AI uses to summarise the page. If they are missing, chance decides. |
| HTTPS | 10 | Unencrypted pages are avoided by crawlers and flagged as insecure by browsers. |
| llms.txt | 10 | The curated content overview for language models, a young standard that signals: this website wants to be understood by AIs. |
| XML sitemap | 10 | The table of contents for crawlers. Without a sitemap, subpages easily go undiscovered. |
| H1 heading | 10 | The main heading is the strongest content signal for what the page is about. |
| Language markup (lang) | 5 | Tells machines which language the page is written in, important for correct attribution in answers. |
| Content Signals | 5 | A young robots.txt standard that declares permitted AI usage. Absence is not an error, setting it is a plus. |
| Extra: Markdown for agents | Bonus | Does the homepage serve a Markdown version on Accept: text/markdown? Saves agents the HTML parsing. |
| Extra: Link header (RFC 8288) | Bonus | Does an HTTP header point agents straight to llms.txt, sitemap or API catalog? |
| Extra: MCP server discovery | Bonus | Is there a server card under /.well-known/mcp/? Own tools for AI assistants, the top tier. |
| Extra: Agent skills | Bonus | Does a skills directory describe ready-made instructions for AI assistants? |
The limits of the scan
The scan reads static HTML. Websites that load their content via JavaScript are seen incompletely, which, to be fair, is exactly how many AI crawlers see them too, so even that is a finding. It checks the homepage plus robots.txt, llms.txt and sitemap, not every subpage.
And the most important limit: technical readability is the prerequisite, not the guarantee. Whether AIs recommend a business is ultimately decided by content, direct answers and authority, which is precisely what GEO consulting works on.
Frequently asked questions
What does the GEO Scan check?
Nine technical foundations of AI readability: HTTPS, how ten AI crawlers are treated in robots.txt, Content Signals, llms.txt, the XML sitemap, Schema.org, title and meta description, language markup and the H1 heading. Plus an extra-credit rating for agent readiness (Markdown negotiation, Link header, MCP server discovery, agent skills), shown as a bonus without lowering the score. Publicly accessible files only, no sign-up.
Why does robots.txt matter so much?
Because many websites lock out AI crawlers without knowing it: common templates and firewall defaults block GPTBot, ClaudeBot or PerplexityBot across the board. A blocked crawler cannot read the site, so the corresponding assistant cannot recommend it.
Does the scan store my data?
No. The scan fetches publicly accessible files, evaluates them on the spot and stores neither the URL nor the result. No sign-up, no cookies.
How is this different from the AI Visibility Check?
The AI Visibility Check measures the outcome: do the AIs know your business? The GEO Scan measures the technical cause: can an AI read your website at all? One, then the other, together they give the full picture.
Scanned red or amber?
Most findings of this scan can be fixed within hours, if you know how. That is exactly the first building block of GEO consulting: repair the foundations, then build content and authority.
Book a first call, free of chargeAlso available to AI assistants
This scan also runs without a browser: it is available as the tool geo_scan on this website's MCP server. AI assistants such as Claude or ChatGPT can scan any website right inside a conversation, using the same logic and the same limits as this page.