| 1 | [cC]laude[bB]ot | claudebot | Anthropic Claude web crawler bot |
| 2 | anthropic-ai | anthropic-ai | Anthropic AI web crawler bot |
| 3 | ApifyBot | Mozilla/5.0 (compatible; ApifyBot/1.0) | Apify's web scraping and data extraction |
| 4 | ApifyWebsiteContentCrawler | ApifyWebsiteContentCrawler/1.0 (+https://apify.com/apify/website-content-crawler) | Apify's full website content extractor |
| 5 | Brightbot | Brightbot 1.0 | Bright Data's AI-ready data collector |
| 6 | Bytespider | Mozilla/5.0 (iPhone; CPU iPhone OS 11_0 like Mac OS X) AppleWebKit/537.36 (KHTML, like Gecko) Chrom… | ByteDance Bytespider web crawler bot |
| 7 | CCBot | CCBot/2.0 (http://commoncrawl.org/faq/) | Common Crawl web crawler for indexing |
| 8 | Channel3Bot | Mozilla/5.0 (compatible; Channel3Bot/1.0; +https://trychannel3.com/channel3bot) | Channel3's universal product catalog indexer |
| 9 | ChatGLM-Spider | Mozilla/5.0 (compatible; ChatGLM-Spider/1.0; +https://chatglm.cn/) | AI model training data collection crawler |
| 10 | ChatGPT-User | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.c… | ChatGPT user web crawler bot |
| 11 | Claude-User | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-User/1.0; +Claude-User@anthro… | Anthropic Claude user web crawler bot |
| 12 | Claude-Web | Claude-Web/1.0 (web crawler; +https://www.anthropic.com/; bots@anthropic.com) | Anthropic Claude web crawler bot |
| 13 | cohere-ai | cohere-ai | Cohere AI web crawler bot |
| 14 | cohere-training-data-crawler | cohere-training-data-crawler (+crawler@cohere.ai) | LLM training data collection for AI models |
| 15 | crawl4ai | crawl4ai-adapter/1.0 | Open-source LLM-friendly web scraping and extraction library |
| 16 | DeepSeekBot | Mozilla/5.0 (compatible; DeepSeekBot/1.0; +https://www.deepseek.com/bot) | DeepSeek AI model training and data collection crawler |
| 17 | ExteContextCrawl | Mozilla/5.0 (compatible; ExteContextCrawl/1.0; +http://crawl001.exte.ai) | AI context extraction and content understanding web crawler |
| 18 | FacebookBot | Mozilla/5.0 (compatible; FacebookBot/1.0; +https://developers.facebook.com/docs/sharing/webmasters/… | Meta's AI speech recognition training crawler |
| 19 | FirecrawlAgent | Mozilla/5.0 (compatible; FirecrawlAgent; +https://firecrawl.dev/) | Firecrawl's LLM data extraction crawler |
| 20 | Flyriverbot | Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.1… | AI content source verification and attribution checking crawler |
| 21 | Google-Extended | Mozilla/5.0 (compatible; Google-Extended/1.0; +http://www.google.com/bot.html) | Google extended web crawler bot |
| 22 | GPTBot | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.0; +https://openai.com/gptb… | OpenAI GPT web crawler bot |
| 23 | HenkBot | Mozilla/5.0 (compatible; HenkBot/1.0; +https://valyu.ai/crawler) | Valyu AI content analysis and data collection crawler |
| 24 | ImageMind | ImageMind | Image analysis crawler collecting visual data for AI training purposes |
| 25 | imageSpider | Mozilla/5.0 (compatible; imageSpider; +https://www.bytedance.com/) | ByteDance's image collection crawler |
| 26 | img2dataset | Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:72.0) Gecko/20100101 Firefox/72.0 (compatible; img2datas… | Open source tool downloading images for machine learning dataset creation |
| 27 | Kangaroo Bot | Mozilla/5.0 (compatible; Kangaroo Bot/1.0; +http://www.kangaroo.com) | Australian AI model training crawler |
| 28 | KunatoCrawler | Mozilla/5.0 (compatible; KunatoCrawler/1.0; +http://kunato.ai/bot.html) | AI training data collection and web scraping crawler |
| 29 | laion-huggingface-processor | Mozilla/5.0 (compatible; laion-huggingface-processor; +https://laion.ai/) | LAION's image dataset builder for AI |
| 30 | MistralAI-User | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-User/1.0; +https://docs.mi… | Mistral AI web crawler bot for content indexing |
| 31 | newsai\/ | Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0… | NewsAI crawler for news content aggregation and indexing |
| 32 | Novellum | Novellum | AI data collection crawler for model training and enrichment |
| 33 | Perplexity-User | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perple… | Perplexity user web crawler bot |
| 34 | PerplexityUser | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityUser/1.0; +https://perplex… | Perplexity AI web crawler bot |
| 35 | Poggio-Citations | Poggio-Citations 1.0 (+https://docs.poggio.io/api/robots) | Poggio's AI sales enablement data collector |
| 36 | SBIntuitionsBot\/ | Mozilla/5.0 (compatible; SBIntuitionsBot/0.1; +https://www.sbintuitions.co.jp/bot/) | SB Intuitions web crawler bot |
| 37 | semantic-visions | Mozilla/5.0 (Linux; CentOS; compatible; semantic-visions-discovery; HTTPClient 4.5) | Semantic content analysis and AI training data crawler |
| 38 | Spawning-AI | Spawning-AI | Web crawler indexing content for AI models |
| 39 | Spider[\s\S]*spider\.com | Mozilla/5.0 (compatible; Spider; +https://www.spider.com/) | AI project web data crawler |
| 40 | TerraCotta | TerraCotta https://github.com/CeramicTeam/CeramicTerracotta | Ceramic Network decentralized data indexing and crawling bot |
| 41 | The Knowledge AI | The Knowledge AI | AI data collection crawler indexing content and enhancing machine learning models |
| 42 | Thinkbot | Mozilla/5.0 (compatible; Thinkbot/0.5.8; +In_the_test_phase,_if_the_Thinkbot_brings_you_trouble,_pl… | Experimental AI thinking and reasoning web crawler bot |
| 43 | TikTokSpider | Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compa… | TikTok web crawler for content discovery |
| 44 | TSM-turingos | Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7; TSM-turingos-1253296984) AppleWebKit/537.36 (KHTML,… | TuringOS AI-powered content analysis and monitoring crawler |