# pietergort.github.io crawl policy # Map: https://pietergort.github.io/llms.txt # Brief: https://pietergort.github.io/machine.md # # Search / user-fetch bots → full site (find + quote). # Training / data-collection bots → machine discovery paths only. User-agent: * Allow: / Sitemap: https://pietergort.github.io/sitemap.xml # --- Search / live citation / user-initiated fetch --- User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / User-agent: PerplexityBot Allow: / # --- Training / foundation-model data collection (machine paths only) --- # Longest-match Allow wins over Disallow: / on major crawlers. User-agent: GPTBot Allow: /machine Allow: /llms Allow: /.well-known/ Disallow: / User-agent: ClaudeBot Allow: /machine Allow: /llms Allow: /.well-known/ Disallow: / User-agent: Google-Extended Allow: /machine Allow: /llms Allow: /.well-known/ Disallow: / User-agent: Applebot-Extended Allow: /machine Allow: /llms Allow: /.well-known/ Disallow: / User-agent: CCBot Allow: /machine Allow: /llms Allow: /.well-known/ Disallow: / User-agent: meta-externalagent Allow: /machine Allow: /llms Allow: /.well-known/ Disallow: / User-agent: Bytespider Allow: /machine Allow: /llms Allow: /.well-known/ Disallow: /