# robots.txt — PMG Engineering # # Stance: open to public crawlers + answer-engine bots. # # We WANT to be cited by ChatGPT / Perplexity / Claude / Google AI # Overviews / Apple Intelligence, so each major AI crawler gets an # explicit Allow block (not just implicit allow-by-omission). This is # the public, machine-readable signal of intent — without it a future # WAF / CDN tweak could silently block one of them. If PMG's policy on # AI training ever changes, flip the relevant `Allow:` to `Disallow:` # here — it's the single canonical source of truth. # # Disallow rules apply to every bot via the shared `*` block at the # bottom: /admin/ is Django admin; /auth/* + /landing/* + /sales/c/* # are SPA shells / token landing pages that have nothing useful for # indexing. # ── AI / LLM crawlers ──────────────────────────────────────────────── # OpenAI — both the training crawler and the runtime fetcher. User-agent: GPTBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: OAI-SearchBot Allow: / # Anthropic — Claude's training + search + runtime crawlers. User-agent: anthropic-ai Allow: / User-agent: ClaudeBot Allow: / User-agent: Claude-Web Allow: / User-agent: Claude-SearchBot Allow: / # Perplexity — both training and runtime. User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # Google's AI-training opt-in (separate from Googlebot). User-agent: Google-Extended Allow: / # Apple's AI-training opt-in (separate from Applebot). User-agent: Applebot-Extended Allow: / # Other notable LLM crawlers (Common Crawl feeds many training sets; # Amazonbot powers Alexa/Q; Bytespider is ByteDance/Doubao). User-agent: CCBot Allow: / User-agent: Amazonbot Allow: / User-agent: Bytespider Allow: / User-agent: Meta-ExternalAgent Allow: / # ── Default rules apply to every crawler (including the AI bots above) ── User-agent: * Disallow: /admin/ # Auth + SPA shells return JS-only HTML; nothing useful to index. Disallow: /auth/ Disallow: /landing/ # Magic-link sales portal token URLs are personalised + ephemeral. Disallow: /sales/c/ # Search results are thin per-query content; the page-level meta # robots already says noindex, but disallowing crawl saves budget. Disallow: /search/ Sitemap: https://pmg.engineering/sitemap.xml