# AI training crawlers — block (no user attached, only feeds model training) User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: Google-Extended Disallow: / User-agent: CCBot Disallow: / User-agent: cohere-ai Disallow: / User-agent: Bytespider Disallow: / User-agent: meta-externalagent Disallow: / User-agent: applebot-extended Disallow: / User-agent: Amazonbot Disallow: / User-agent: Diffbot Disallow: / User-agent: FacebookBot Disallow: / User-agent: Omgilibot Disallow: / User-agent: PetalBot Disallow: / # SEO scrapers — block (high-volume, no SEO value to us) User-agent: DataForSeoBot Disallow: / User-agent: SemrushBot Disallow: / User-agent: AhrefsBot Disallow: / User-agent: MJ12bot Disallow: / User-agent: DotBot Disallow: / # AI user-driven retrieval — allow (these fetch on behalf of a real user query) User-agent: ChatGPT-User Allow: / User-agent: OAI-SearchBot Allow: / User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / User-agent: Claude-Web Allow: / User-agent: Applebot Allow: / # Default policy — index the catalog, skip action endpoints and helpers User-agent: * Disallow: /api/ Disallow: /logos/ Disallow: /rt/ Disallow: /rp Disallow: /cr/ Disallow: /np/ Disallow: /su/ Disallow: /altcha.php Disallow: /load-more.php Disallow: /analytics.php Disallow: /nowplaying.php Disallow: /rating.php Disallow: /report.php Disallow: /check-report.php Disallow: /stats Disallow: /health Disallow: /assets/ # Content usage signal (Cloudflare Content-Signal): search indexing and # user-driven AI retrieval are welcome (they send real visitors back); using # the catalog for AI model training is not. Content-Signal: search=yes, ai-input=yes, ai-train=no Sitemap: https://streamurl.link/sitemap.xml