# Walkie Talkie — robots.txt # # Standard search-engine crawlers are welcome — POI / city / venue # pages benefit from being indexed by Google, Bing, DuckDuckGo, etc. # That's the SEO moat. # # AI training crawlers are NOT welcome. The curated venue content # (theme parks, zoos, exhibitions, highlights, route prompts) is # the result of significant human + editorial work and we don't # consent to it being used to train commercial LLMs without a # licensing agreement. # # This file signals intent. Well-behaved bots respect it; bad # actors don't. For the bad-actor case server-side rate limiting + # IP throttling is the second layer (see server/middleware). # ─── Allow normal search indexing ──────────────────────────────── User-agent: Googlebot Allow: / User-agent: Bingbot Allow: / User-agent: DuckDuckBot Allow: / User-agent: Applebot Allow: / User-agent: Slurp Allow: / # ─── Block AI training crawlers ────────────────────────────────── # OpenAI User-agent: GPTBot Disallow: / User-agent: ChatGPT-User Disallow: / User-agent: OAI-SearchBot Disallow: / # Anthropic User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: Claude-Web Disallow: / # Google's "extended" crawler — separate opt-out from Googlebot so # you can keep search indexing while denying Gemini training data. User-agent: Google-Extended Disallow: / # Common Crawl — feeds many other LLMs downstream User-agent: CCBot Disallow: / # Perplexity User-agent: PerplexityBot Disallow: / # ByteDance / TikTok User-agent: Bytespider Disallow: / # Meta / Facebook AI User-agent: FacebookBot Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: meta-externalagent Disallow: / # Amazon User-agent: Amazonbot Disallow: / # Apple (general LLM training crawler, distinct from Applebot for # Siri/Spotlight which IS allowed above) User-agent: Applebot-Extended Disallow: / # Cohere User-agent: cohere-ai Disallow: / # Other AI scrapers User-agent: Diffbot Disallow: / User-agent: Omgili Disallow: / User-agent: omgilibot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: Timpibot Disallow: / User-agent: VelenPublicWebCrawler Disallow: / # ─── Default ────────────────────────────────────────────────────── # Everything else not explicitly listed above is allowed by default. # Adjust if you want a stricter posture. # Sitemap. The prerender step (scripts/prerender.mjs) replaces this # placeholder line in dist/robots.txt with the real Sitemap: URL, # stamped from SITE_URL. Keep the leading "# " — it's the regex # anchor the replacement matches on. Sitemap: https://walkietalkie.guide/sitemap.xml