User-agent: * Disallow: /admin/ Disallow: /cart Disallow: /*?sessionid= Allow: /admin/public/ User-agent: GPTBot User-agent: ClaudeBot User-agent: CCBot Disallow: / Sitemap: https://example.com/sitemap.xml
Upload this to the root of your domain so it is served at /robots.txt.
The longest matching path pattern wins, and Allow beats Disallow on a tie - the same precedence real crawlers use.
Click Block to add a group that refuses a crawler outright. Compliance is voluntary - this stops cooperative bots, not scrapers.
| Token | What it is | Action |
|---|---|---|
| Googlebot | Google Search's main web crawler. Blocking it removes you from Google over time. | |
| Googlebot-Image | Crawls images for Google Images. Block it to keep photos out of image search. | |
| Bingbot | Microsoft Bing's crawler. Also feeds Bing-powered results in other products. | |
| DuckDuckBot | DuckDuckGo's own crawler for its instant answers and index. | |
| Baiduspider | Baidu's crawler, the dominant search engine in China. | |
| YandexBot | Yandex's crawler, the leading search engine in Russia. | |
| Slurp | Yahoo's legacy crawler. Still used for some Yahoo verticals. | |
| Applebot | Apple's crawler, powering Siri and Spotlight Suggestions. | |
| Twitterbot | Fetches link preview cards for X/Twitter. Blocking it kills your share previews. | |
| facebookexternalhit | Fetches Open Graph previews for Facebook, Messenger, and WhatsApp. | |
| GPTBotAI | OpenAI's crawler that collects public web content used to train its models. | |
| ClaudeBotAI | Anthropic's crawler for gathering training data for Claude. | |
| PerplexityBotAI | Perplexity's crawler, used to index pages its answer engine cites. | |
| CCBotAI | Common Crawl's bot. Its public corpus is a common source of AI training data. | |
| Google-ExtendedAI | Controls whether Google may use your content for Gemini and Vertex AI training. It does not affect Google Search ranking or indexing. | |
| BytespiderAI | ByteDance's crawler, associated with TikTok and its AI products. | |
| anthropic-aiAI | Legacy Anthropic token. Harmless to keep alongside ClaudeBot for older references. |

robots.txt 產生器建構位於您網域根目錄的純文字檔案,並告訴爬蟲他們可能取得的網站的哪些部分。 格式看起來很簡單 - 一個 `user-agent:` 行,後面跟著 `allow:` 和 `disallow:` 規則 - 但細節 bit: 規則只適用於群組,路徑必須以斜線、`*` 和 `$` 以特定方式開始,並且在 `user-agent: *` 下可以將整個網站拉出 Google 之外。 此工具從結構化表單建構檔案,因此語法透過結構進行正確。
它做了三項以外的工作。 它掛牌並標記真正的足球:規則漂浮在任何群組之外,相對路徑,一個裸露的“不允許:”靜默意味著“允許一切”,不是絕對的網站地圖 URL,未知指令,以及毯子塊。 它將現有的 robots.txt 解析為可編輯的群組,因此您可以貼上今天的即時內容並修復它。 它使用與真實爬蟲相同的最長優先級來測試您的規則 - 最具體的匹配模式獲勝,並允許在 tie 上禁止使用節拍 - 這樣您就可以在部署之前證明頁面是可訪問的。
值得簡單陳述的一件事:robots.txt 控制爬行,而不是索引,它不是安全邊界。 如果其他頁面連結到它,則阻止的 URL 仍然可以出現在搜尋結果中,因為 Google 可以索引它從未取得的 URL。 要將頁面排除在索引之外,您需要一個 `noindex` 元標記或 X-robots-tag 標題 - 爬蟲只能查看是否允許取出頁面。 該文件是公開的:任何人都可以在 /robots.txt 上閱讀您的文件,因此在那裡列出秘密目錄會做廣告。 對任何私有的使用身份驗證。
選擇預設 - 允許一切,阻止一切,WordPress,Shopify,Next.js 或 Block AI 爬蟲 - 一鍵填寫表單,或粘貼您的 Live Robots.txt 並讓該工具將其解析回可編輯的群組。
每個爬蟲或一組爬蟲添加一個群組,從已知的機器人清單中選擇或鍵入一個令牌。 給每個組別不允許和允許路徑,如果您瞄準的引擎 Honour 一號,則為 Crawl-Delay 提供。
新增一個或多個絕對網站地圖 URL。 驗證器在每次按鍵上運行並報告錯誤、警告和帶有行號的註釋 - 在您發貨之前修復任何紅色。
輸入路徑和使用者代理以查看允許或不允許的規則以及確切的規則決定了它。 當它按照您的預期方式行事時,請複製檔案或下載 robots.txt 並將其上傳到您的網域根。
使用 RFC 9309 優先權對任何用戶代理測試任何路徑 - 最長的匹配模式獲勝,允許斷線 - 並命名決定它的確切規則
在群組之外捕捉規則、相對路徑、空的不允許行、非絕對網站地圖 urls、未知指令和一攬子不允許:/這將取消索引您的網站
允許一切,阻止一切,wordpress,shopify,next.js 和 block ai crawlers 每個都用一個明智的、正確的起點填滿整個表單
Block GPTBot、Claudebot、PerplexityBot、CCBot、Google 擴展、Bytespider 和 Anthropic-AI,同時讓搜尋引擎充分歡迎
Free hreflang generator for multilingual and multi-region sites - build hreflang tags as HTML link tags, XML sitemap entries, or an HTTP header. Validates codes and x-default. 100% client-side.
Generate Apache .htaccess 301 and 302 redirect rules from a list of old-to-new URLs, plus force-HTTPS and www canonicalization directives, all in the browser.
Check one name across hundreds of TLDs at once. RDAP first, WHOIS fallback, and a verdict of available, taken or unknown — never a guess
Pull business listings for any search and place: name, address, category, rating, review count, phone and website, deduplicated and exportable as CSV.
0 comments
No comments yet. Be the first to share your thoughts!