User-agent: * Disallow: /admin/ Disallow: /cart Disallow: /*?sessionid= Allow: /admin/public/ User-agent: GPTBot User-agent: ClaudeBot User-agent: CCBot Disallow: / Sitemap: https://example.com/sitemap.xml
Upload this to the root of your domain so it is served at /robots.txt.
The longest matching path pattern wins, and Allow beats Disallow on a tie - the same precedence real crawlers use.
Click Block to add a group that refuses a crawler outright. Compliance is voluntary - this stops cooperative bots, not scrapers.
| Token | What it is | Action |
|---|---|---|
| Googlebot | Google Search's main web crawler. Blocking it removes you from Google over time. | |
| Googlebot-Image | Crawls images for Google Images. Block it to keep photos out of image search. | |
| Bingbot | Microsoft Bing's crawler. Also feeds Bing-powered results in other products. | |
| DuckDuckBot | DuckDuckGo's own crawler for its instant answers and index. | |
| Baiduspider | Baidu's crawler, the dominant search engine in China. | |
| YandexBot | Yandex's crawler, the leading search engine in Russia. | |
| Slurp | Yahoo's legacy crawler. Still used for some Yahoo verticals. | |
| Applebot | Apple's crawler, powering Siri and Spotlight Suggestions. | |
| Twitterbot | Fetches link preview cards for X/Twitter. Blocking it kills your share previews. | |
| facebookexternalhit | Fetches Open Graph previews for Facebook, Messenger, and WhatsApp. | |
| GPTBotAI | OpenAI's crawler that collects public web content used to train its models. | |
| ClaudeBotAI | Anthropic's crawler for gathering training data for Claude. | |
| PerplexityBotAI | Perplexity's crawler, used to index pages its answer engine cites. | |
| CCBotAI | Common Crawl's bot. Its public corpus is a common source of AI training data. | |
| Google-ExtendedAI | Controls whether Google may use your content for Gemini and Vertex AI training. It does not affect Google Search ranking or indexing. | |
| BytespiderAI | ByteDance's crawler, associated with TikTok and its AI products. | |
| anthropic-aiAI | Legacy Anthropic token. Harmless to keep alongside ClaudeBot for older references. |

robots.txt 生成器构建位于您的域根的纯文本文件,并告诉爬虫他们可能获取的网站的哪些部分。 格式看起来很简单 - 一个 `user-agent:` 行,后面跟着`allow:` 和 `disallow:` 规则——但细节咬合:规则只适用于组内,路径必须以斜杠、`*` 和 `$` 以特定方式开始,并且单一的流浪 `disallow: /` 下 `user-agent: ` 下的 /` 下 `user-agent: ` 可以将整个站点从 Google 中拉出。 该工具从结构化表单构建文件,因此语法通过构造是正确的。
它完成了三项超出世代的工作。 它会将文件标记为真实的行列:浮动的规则,在任何组之外,相对路径,一个裸露的`disallow:`,它默默地表示“允许一切”,站点地图的URL不是绝对的,未知的指令和一揽子块。 它将现有的 robots.txt 解析为可编辑的组,因此您可以粘贴今天的实时内容并进行修复。 它会根据您的规则与真实爬虫使用的最长匹配优先级测试您的规则 - 最具体的匹配模式获胜,并允许在 TIE 上不允许节拍 - 这样您就可以在部署之前证明页面是可访问的。
值得一提的是: robots.txt 控制爬行,而不是索引,它不是安全边界。 如果其他网页链接到该网页,则仍会显示阻止的网址,因为 Google 可以将其从未获取的网址编入索引。 要将页面排除在索引之外,您需要一个 `noindex` 元标记或一个 x-robots-tag 标头 - 爬网程序只能看到它是否被允许获取页面。 并且该文件是公开的:任何人都可以在 /robots.txt 中读取您的文件,因此列出了一个秘密目录在其中进行广告。 对任何私有的东西使用身份验证。
选择一个预设 - 允许所有内容,阻止所有内容,WordPress,Shopify,Next.js 或 Block AI 爬虫 - 一键填写表单,或粘贴您的实时 robots.txt 并让工具将其解析回可编辑的组中。
每个抓取工具或一组爬网器添加一个组,从已知的机器人列表中进行选择或键入令牌。 给每组它的禁止和允许路径,如果你的引擎是你所针对的引擎,则进行爬行延迟。
添加一个或多个绝对站点地图 URL。 验证器在每次击键上运行并报告错误、警告和带有行号的注释 - 修复在发货前的红色任何内容。
输入路径和用户代理以查看允许或不允许以及确切的规则决定了它。 当它按照您期望的方式运行时,请复制文件或下载 robots.txt 并将其上传到您的域根。
使用RFC 9309优先级测试任何用户代理的路径 - 最长的匹配模式获胜,允许断开联系 - 并命名决定它的确切规则
捕获组之外的规则、相对路径、空的不允许行、非绝对站点地图 URL、未知指令和一揽子不允许:/ 这将取消索引您的网站
允许所有内容,阻止所有内容,WordPress,Shopify,Next.js和Block AI爬虫,每个人都用一个明智的、正确的起点填满整个表格
一键阻止GPTBOT,Claudebot,PleprepleityBot,CCBOT,Google-Extended,Bytespider和Anthropic-AI,同时保持搜索引擎的全力欢迎
Free hreflang generator for multilingual and multi-region sites - build hreflang tags as HTML link tags, XML sitemap entries, or an HTTP header. Validates codes and x-default. 100% client-side.
Generate Apache .htaccess 301 and 302 redirect rules from a list of old-to-new URLs, plus force-HTTPS and www canonicalization directives, all in the browser.
Check one name across hundreds of TLDs at once. RDAP first, WHOIS fallback, and a verdict of available, taken or unknown — never a guess
Pull business listings for any search and place: name, address, category, rating, review count, phone and website, deduplicated and exportable as CSV.
0 comments
No comments yet. Be the first to share your thoughts!