一、采集脚本 / 命令行工具(83 条)
当前页面展示的爬虫,视业务自行取舍;自写脚本与无浏览器特征的 HTTP 客户端,是访问日志中"非正常流量"的最大来源。特征:UA 与真实浏览器毫无相似性,极易识别。
表格中的:"官方说明"一列保留英文原文描述,避免翻译引入歧义;"匹配正则"可直接用于做访问日志分析或 WAF 拦截规则。
1.1 命令行工具(CLI)
| # | 匹配正则 | 典型 UA 实例 | 官方说明 |
|---|---|---|---|
| 1 | [wW]get | WGETbot/1.0 (+http://wget.alanreed.org) | GNU Wget command-line tool for downloading web content |
| 2 | ^curl | curl | cURL command-line tool for HTTP requests |
| 3 | ^HTTPie\/ | HTTPie/3.2.2 | HTTPie command-line HTTP client tool |
| 4 | http_get | http_get | HTTP GET command-line tool for requests |
1.2 HTTP 客户端库 / 编程语言 SDK
| # | 匹配正则 | 典型 UA 实例 | 官方说明 |
|---|---|---|---|
| 1 | [mM]echanize | Mechanize/2.9.1 Ruby/3.1.2 (http://github.com/sparklemotion/mechanize/) | Mechanize Ruby web crawler library |
| 2 | ^Apache-HttpClient | Apache-HttpClient/4.2.3 (java 1.5) | Apache HTTP Client Java library for requests |
| 3 | ^PHP-Curl-Class | PHP-Curl-Class/4.13.0 (+https://github.com/php-curl-class/php-curl-class) PHP/7.2.24 curl/7.61.1 | PHP Curl Class HTTP client library |
| 4 | AHC\/ | AHC/2.0 | Async HTTP Client Java library for HTTP requests |
| 5 | aiohttp | Python/3.9 aiohttp/3.7.3 | Asynchronous HTTP client library for Python |
| 6 | ALittle Client | ALittle Client | ALittle web crawler for content discovery |
| 7 | Amazon CloudFront | Amazon CloudFront | Amazon CloudFront web crawler bot |
| 8 | AnyEvent | Mozilla/5.0 (compatible; U; AnyEvent-HTTP/2.24; +http://software.schmorp.de/pkg/AnyEvent) | AnyEvent Perl HTTP web crawler library |
| 9 | axios | axios/0.18.0 | Axios HTTP client library for requests |
| 10 | BTWebClient | BTWebClient/180B(9704) | µTorrent BitTorrent Client |
| 11 | colly | colly - https://github.com/gocolly/colly | Colly Go web crawler framework |
| 12 | crawler4j | crawler4j (http://code.google.com/p/crawler4j/) | Crawler4j Java web crawler framework |
| 13 | CustomAsyncHttpClient | CustomAsyncHttpClient | Custom async HTTP client web crawler |
| 14 | Download Ninja | Download Ninja/4.0 | File download manager and web content fetching tool |
| 15 | Go-http-client | Go-http-client/1.1 | Go programming language HTTP client library |
| 16 | GoParserBot | Mozilla/5.0 (compatible; GoParserBot/1.0) | Go-based web content parsing and extraction crawler bot |
| 17 | httpunit | httpunit/1.x | Java library for automated web application testing |
| 18 | HttpUrlConnection | Jersey/2.25.1 (HttpUrlConnection 1.8.0_141) | Java HttpUrlConnection HTTP client |
| 19 | httpx | python-httpx/0.16.1 | Modern Python HTTP client with async support |
| 20 | Jetty | Jetty/9.3.z-SNAPSHOT | Jetty web server HTTP client library |
| 21 | larbin | larbin_2.6.2 (larbin@correa.org) | Open-source web crawler for large-scale indexing projects |
| 22 | libwww-perl | 2Bone_LinkChecker/1.0 libwww-perl/6.03 | Perl library for making HTTP requests and web crawling |
| 23 | lwp-trivial | lwp-trivial/1.35 | Perl LWP library simple HTTP client and fetcher |
| 24 | newspaper\/ | newspaper/0.1.0.7 | Newspaper Python web scraping library |
| 25 | node-fetch | node-fetch/1.0 (+https://github.com/bitinn/node-fetch) | Node-fetch HTTP client library |
| 26 | okhttp | okhttp/2.5.0 | OkHttp Java HTTP client library |
| 27 | Pcore-HTTP | Pcore-HTTP/v0.40.3 | Pcore HTTP web crawler library |
| 28 | phpcrawl | phpcrawl | PHP web crawler library for scraping websites |
| 29 | python-opengraph | python-opengraph-jaywink/0.2.0 (+https://github.com/jaywink/python-opengraph) | Python OpenGraph web crawler bot |
| 30 | python-requests | python-requests/2.9.2 | Popular Python HTTP library for making web requests |
| 31 | Python-urllib | Python-urllib/1.17 | Python's built-in URL library for HTTP requests |
| 32 | rawweb-bot | rawweb-bot/1.0 | Raw web content extraction and data scraping crawler |
| 33 | Salesforce\.com | Salesforce.com | Salesforce HTTP callout user agent |
| 34 | Scrapy | Scrapy/1.0.3 (+http://scrapy.org) | Scrapy Python web crawler framework |
| 35 | SimpleCrawler | SimpleCrawler/0.1 | Simple Crawler web crawler framework |
| 36 | sindresorhus\/got | got (https://github.com/sindresorhus/got) | Got HTTP client library for requests |
| 37 | trafilatura | trafilatura/2.0.0 (+https://github.com/adbar/trafilatura) | Python library for web content extraction and scraping |
1.3 采集框架 / 脚本工具
| # | 匹配正则 | 典型 UA 实例 | 官方说明 |
|---|---|---|---|
| 1 | AdminLabs | AdminLabs | Web scraping tool for data collection and analysis |
| 2 | allOrigins | Mozilla/5.0 (compatible; allOrigins/3; +http://allorigins.win/) | Web scraping proxy tool bypassing CORS restrictions |
| 3 | CyotekWebCopy | CyotekWebCopy/1.9 CyotekHTTP/6.4 | Website copying and offline browsing tool crawler bot |
| 4 | eMoneyBot | eMoneyBot/1.0; +https://emoneyadvisor.com/DataAggregationNotice/ | Financial account aggregation and data scraping |
| 5 | EpivozCrawler | EpivozCrawler/1.7 | Content scraping and indexing crawler |
| 6 | ExodusMovement | ExodusMovement/1.0 GlobalCoinHeight/1.0 | Web scraping and data extraction bot |
| 7 | Grover\/ | Grover/Grover-1.18 (Web Crawler) | Anonymous web crawler likely used for research or content scraping |
| 8 | HTTrack | Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98) | HTTrack website copier for offline browsing |
| 9 | KangarooBot\/ | Mozilla/5.0 (compatible; KangarooBot/1.0; +https://kangaroobot.com) | Data scraping and content aggregation crawler mimicking human behavior |
| 10 | KrawlerBot | Mozilla/5.0 (compatible; KrawlerBot; +https://www.krawler.com/) | Krawler web scraping service bot |
| 11 | noorobot | noorobot | Web scraper performing high-volume programmatic content collection |
| 12 | Offline Explorer | Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; Offline Explorer; MSIECrawler) | Website downloader tool for offline archival and content analysis |
| 13 | s4a-probe-bot\/ | Mozilla/5.0 AppleWebKit (compatible; s4a-probe-bot/1.0; Fake-Googlebot; +https://www.seo4ajax.com/w… | Web scraping and SEO analysis crawler tool |
| 14 | SBL-BOT | SBL-BOT (http://sbl.net) | Bot of SoftByte BlackWidow |
| 15 | ScrapeheroBot\/ | Mozilla/5.0 (compatible; ScrapeheroBot/1.0; +https://scrapehero.de/) | Web scraping service crawler for data extraction |
| 16 | SimpleScraper | Mozilla/5.0 (compatible; SimpleScraper) | Simple Scraper web scraping tool |
| 17 | SiteSucker | SiteSucker for macOS/6.1.5 | macOS website downloading and offline browsing tool crawler |
| 18 | Trellis-Services | Trellis-Services | Content scraper and fetching service crawler for programmatic large-scale web page retrieval |
| 19 | WebCopier | WebCopier v4.6 | Website copying and offline browsing tool crawler bot |
| 20 | WebZIP | WebZIP/3.5 (http://www.spidersoft.com) | Website downloading and offline browsing tool crawler bot |
| 21 | WSM\/ | Mozilla/5.0 (compatible; WSM/2.0; +https://webspidermount.com/) | Web scraper crawler for content collection and data aggregation purposes |
1.4 浏览器自动化 / 无头浏览器
| # | 匹配正则 | 典型 UA 实例 | 官方说明 |
|---|---|---|---|
| 1 | Amazon-Bedrock-AgentCore-Browser | Mozilla/5.0 (compatible; Amazon-Bedrock-AgentCore-Browser; +https://aws.amazon.com/bedrock/) | AWS cloud browser for AI agents |
| 2 | AmazonBuyForMe | Mozilla/5.0 (compatible; AmazonBuyForMe; +https://www.amazon.com/) | Amazon bot for Buy For Me service purchases |
| 3 | Anchor Browser | Mozilla/5.0 (compatible; Anchor Browser; +https://anchorbrowser.io/) | Anchor's cloud-hosted browser for AI agents |
| 4 | Code\/1\. | Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Code/1.115.0 Chrom… | GitHub Copilot AI coding agent |
| 5 | Devin | Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/137.0.0.0 Safari/537.… | AI-driven browser automation and content extraction |
| 6 | Fluid | Mozilla/5.0 (Macintosh; U; Intel Mac OS X 10_5_6; en-us) AppleWebKit/528.16 (KHTML, like Gecko) Flu… | Site-specific browser application for web content access |
| 7 | Ghost Inspector | Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/107.0 Safari/537.36 G… | Browser testing service automating UI tests and monitoring website changes |
| 8 | Google-Agent | Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko; compatible; Google-Agent) | Google's user-triggered agent for web navigation |
| 9 | Google-Gemini-CLI | Mozilla/5.0 (compatible; Google-Gemini-CLI/1.0; +https://github.com/google-gemini/gemini-cli) | Google's AI coding agent for terminal |
| 10 | GoogleAgent-Mariner | Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko; compatible; GoogleAgent-Mari… | Google's AI agent for browser-based tasks |
| 11 | HeadlessChrome | Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) HeadlessChrome/74.0.3729.169… | Headless Chrome web crawler for testing |
| 12 | Manus-User | Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/132.0… | Butterfly Effect's AI agent for web tasks |
| 13 | opencode-smartfetch | opencode-smartfetch/1.0 | Open-source AI coding agent |
| 14 | PhantomJS | Mozilla/5.0 (Unknown; Linux x86_64) AppleWebKit/538.1 (KHTML, like Gecko) PhantomJS/2.1.1 Safari/53… | PhantomJS headless browser web crawler |
| 15 | Playwright | Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 S… | Playwright browser automation web crawler |
| 16 | Puppeteer | Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) HeadlessChrome/120.0.0.0 Saf… | Puppeteer browser automation web crawler |
| 17 | Selenium | Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 S… | Selenium browser automation web crawler |
| 18 | splash Version\/ | Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/602.1 (KHTML, like Gecko) splash Version/10.0 Safari/60… | Headless browser for web scraping and testing |
| 19 | TestLocally\/ | Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.6422.26 Safari/… | Development testing bot simulating user interactions and verifying website functionality behavior |
| 20 | Trae\/ | Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Trae/1.107.1 Chrom… | ByteDance's AI coding agent |
| 21 | TwinAgent | Mozilla/5.0 (compatible; TwinAgent; +https://www.twinagent.com/) | Twin's automated worker for workflow execution |
