Robots · TXT # robots.txt # Blocks known AI/training crawlers, aggressive SEO scrapers, and generic # scraping tools, while leaving mainstream search engines free to index. # # IMPORTANT: robots.txt is a voluntary honor-system file. Well-behaved bots # (Google, Bing, the ones listed below) respect it. Malicious bots and # scrapers that are actually trying to hide will simply ignore it — this # file is not a security control. If you need real enforcement, block by # user-agent/IP at your firewall, CDN, or reverse proxy (e.g. Cloudflare # Bot Fight Mode, or an nginx/Apache user-agent rule) in addition to this. # --------------------------------------------------------------------- # AI crawlers / LLM training & retrieval bots # --------------------------------------------------------------------- User-agent: CCBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: GoogleOther Disallow: / User-agent: Bytespider Disallow: / User-agent: PerplexityBot Disallow: / User-agent: Perplexity-User Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Amazonbot Disallow: / User-agent: cohere-ai Disallow: / User-agent: Diffbot Disallow: / User-agent: FacebookBot Disallow: / User-agent: meta-externalagent Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: Timpibot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: Omgilibot Disallow: / User-agent: Omgili Disallow: / User-agent: YouBot Disallow: / # --------------------------------------------------------------------- # Aggressive SEO / marketing crawlers (heavy load, little user value) # --------------------------------------------------------------------- User-agent: MJ12bot Disallow: / User-agent: DotBot Disallow: / User-agent: BLEXBot Disallow: / User-agent: DataForSeoBot Disallow: / User-agent: SerpstatBot Disallow: / User-agent: Barkrowler Disallow: / User-agent: Seekr Disallow: / # --------------------------------------------------------------------- # Generic scraping libraries / tools (default user-agent strings) # --------------------------------------------------------------------- User-agent: python-requests Disallow: / User-agent: Scrapy Disallow: / User-agent: curl Disallow: / User-agent: wget Disallow: / User-agent: HeadlessChrome Disallow: / # --------------------------------------------------------------------- # Everyone else (mainstream search engines, etc.) — allowed # --------------------------------------------------------------------- User-agent: * Allow: / Sitemap: https://www.wordsatwork.com/sitemap.xml Sitemap: https://www.wordsatwork.com/sitemap.xml