Checked 9/4/2026 · https://example.org/
No hard blockers, but 4 points weaken AI discoverability.
AI crawler matrix
For each bot: what your robots.txt allows — and what your server actually returns to the bot's user agent. Rows highlighted in red: robots allows it, but the server blocks it (usually a WAF or bot-protection rule).
| Bot | Operator | robots.txt | Server response |
|---|---|---|---|
| GPTBot | OpenAI | not named | 200 + content |
| OAI-SearchBot | OpenAI | not named | 200 + content |
| ChatGPT-User | OpenAI | not named | 200 + content |
| ClaudeBot | Anthropic | not named | 200 + content |
| Claude-SearchBot | Anthropic | not named | 200 + content |
| Claude-User | Anthropic | not named | 200 + content |
| PerplexityBot | Perplexity | not named | 200 + content |
| Perplexity-User | Perplexity | not named | 200 + content |
| Google-Extended | Google (Gemini-Training) | not named | — |
| Googlebot | Google (KI-Übersichten) | not named | 200 + content |
| CCBot | Common Crawl (Trainingsdaten vieler LLMs) | not named | 200 + content |
| Bytespider | ByteDance (Doubao) | not named | 200 + content |
| Amazonbot | Amazon (Alexa/Rufus) | not named | 200 + content |
| Applebot-Extended | Apple (Intelligence-Training) | not named | — |
| Meta-ExternalAgent | Meta (Llama/Meta AI) | not named | 200 + content |
How is the score calculated?
| Very little text in the HTML | -10 |
| No structured data (JSON-LD) | -8 |
| Page title too short or too long | -6 |
| Meta description missing or unsuitable | -5 |
| No llms.txt | -5 |
| No author or authorship signal | -4 |
| No sitemap.xml found | -4 |
| No sameAs links | -3 |
| No link to legal notice / about / contact | -3 |
| No visible or marked-up date | -3 |
| No canonical URL | -3 |
| No robots.txt found | -2 |
| Result | 44 / 100 |
Access for AI crawlers
No robots.txt found
No valid file was reachable at /robots.txt. Crawlers then assume "everything allowed" — it works, but you have no control and no sitemap signal.
This is what you should do: Create a robots.txt in the root directory, at least with a sitemap reference and your desired allow/disallow rules.
User-agent: * Allow: / Sitemap: https://your-domain.com/sitemap.xml
Content without JavaScript
Very little text in the HTML
The initial HTML contains only around 142 characters of visible text. That's barely enough for AI systems to grasp the page's topic with confidence.
This is what you should do: Make sure the actual content is fully in the HTML and not in images, iframes or blocks loaded via JavaScript.
Structure & machine understanding
No structured data (JSON-LD)
The page contains no JSON-LD. Structured data helps AI systems recognise entities, author, organisation, date and page type unambiguously — without it, everything has to be guessed from the body text.
This is what you should do: Add at least an "Organization" schema and one that fits the page type (Article, Product, FAQPage …) as JSON-LD in the <head>.
<script type="application/ld+json">
{"@context":"https://schema.org","@type":"Organization",
"name":"Your company","url":"https://your-domain.com",
"sameAs":["https://www.linkedin.com/company/…"]}
</script>Page title too short or too long
The <title> is 14 characters long. Around 30–60 characters is sensible: descriptive enough to name the topic, short enough not to be cut off.
This is what you should do: Write a concise title that contains the topic and, if relevant, the brand.
Meta description missing or unsuitable
The meta description is 0 characters long. It is often the first thing AI systems and search engines read as a summary of the page.
This is what you should do: Write a standalone summary of the page in one or two sentences (around 120–200 characters), not a string of keywords.
No llms.txt
There is no file at /llms.txt. llms.txt is a short Markdown overview of your key content that some AI systems use for orientation.
This is what you should do: Create an /llms.txt: an H1 with the name, a short paragraph about the offering, and a link list to the central pages.
# Your company > One-sentence description. ## Key pages - [Product](https://your-domain.com/product): … - [Pricing](https://your-domain.com/pricing): … - [About](https://your-domain.com/about): …
Trust and entity signals (E-E-A-T)
No author or authorship signal
No author is recognisable (neither meta tag, rel=author nor JSON-LD Person). For "experience" and "expertise" in the E-E-A-T sense, named, traceable authorship matters.
This is what you should do: For editorial content, name an author and link an author page; also mark them up as JSON-LD "Person".
No sameAs links
The JSON-LD is missing "sameAs". It lets AI systems clearly map your brand or person to known entities (Wikidata, LinkedIn, industry directories).
This is what you should do: Add a "sameAs" in Organization or Person with the URLs of your official profiles and directory entries.
No link to legal notice / about / contact
No link to a legal notice, "About" or contact page is recognisable on the checked page. Such pages are a strong trust signal and help with entity mapping.
This is what you should do: Link the legal notice or "About" and contact clearly, usually in the footer of every page.
No visible or marked-up date
No publication or modification date is recognisable (neither <time>, article:published_time nor datePublished in the JSON-LD). For many questions, AI systems prefer current sources.
This is what you should do: For content with a time reference, show a date visibly and mark up datePublished / dateModified in the JSON-LD.
Technical foundation
No sitemap.xml found
No valid XML sitemap was reachable at /sitemap.xml and robots.txt names none. A sitemap helps crawlers find all relevant URLs.
This is what you should do: Generate a sitemap.xml with all indexable URLs and reference it in robots.txt.
No canonical URL
The page sets no <link rel="canonical">. Without a canonical, it can stay unclear which version is authoritative when several URL variants exist.
This is what you should do: Set a self-referencing <link rel="canonical"> with the preferred URL on every page.
OK (2)
- Exactly one H1 heading.
- Fast server response time.
Next steps
- Very little text in the HTML. Make sure the actual content is fully in the HTML and not in images, iframes or blocks loaded via JavaScript.
- No structured data (JSON-LD). Add at least an "Organization" schema and one that fits the page type (Article, Product, FAQPage …) as JSON-LD in the <head>.
- Page title too short or too long. Write a concise title that contains the topic and, if relevant, the brand.
Weekly re-scan with an email alert on every change. In preparation — add your address.
Method & limits
Checked 9/4/2026. citeglass fetches example.org and its related files (robots.txt, llms.txt, sitemap.xml) over HTTP — once as a normal browser, once per AI crawler user agent. No JavaScript is executed. Time to first byte: 21 ms.
What this is not: No rank or citation tracking, no statement about whether a model actually names you, and no check of content loaded via JavaScript. A snapshot from the perspective of one server IP.