User-agent: Googlebot Allow: / User-agent: Bingbot Allow: / User-agent: Twitterbot Allow: / User-agent: facebookexternalhit Allow: / # Everything below is public marketing content: no login, no paywall, no # customer data. There is nothing here worth hiding from an AI crawler, and # real upside in being read and cited correctly, so these are named # explicitly and allowed rather than left to fall through to the wildcard. # OpenAI: GPTBot trains on and indexes content, ChatGPT-User fetches a page # a person asked ChatGPT to open, OAI-SearchBot powers ChatGPT's web search # and citations. User-agent: GPTBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: OAI-SearchBot Allow: / # Anthropic: ClaudeBot crawls and indexes, Claude-User fetches a page a # person asked Claude to open (as this task is doing right now), Claude- # SearchBot powers Claude's web search and citations. User-agent: ClaudeBot Allow: / User-agent: Claude-User Allow: / User-agent: Claude-SearchBot Allow: / # Perplexity: PerplexityBot builds its search index, Perplexity-User fetches # a page on a person's behalf. User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # Google-Extended and Applebot-Extended are separate opt-ins from Googlebot # and Applebot, covering use in Gemini/AI Overviews and Apple Intelligence. User-agent: Google-Extended Allow: / User-agent: Applebot-Extended Allow: / # Meta's AI crawler (distinct from facebookexternalhit's link-preview bot) # and Amazon's, which underlies Alexa+ answers. User-agent: meta-externalagent Allow: / User-agent: Amazonbot Allow: / User-agent: * Allow: / Sitemap: https://theautomate.io/sitemap.xml # /llms.txt describes this same site in plain language for language models: # what the agency does, and a line on every page.