User-agent: * Disallow: /admin/ Disallow: /cart Disallow: /*?sessionid= Allow: /admin/public/ User-agent: GPTBot User-agent: ClaudeBot User-agent: CCBot Disallow: / Sitemap: https://example.com/sitemap.xml
Upload this to the root of your domain so it is served at /robots.txt.
The longest matching path pattern wins, and Allow beats Disallow on a tie - the same precedence real crawlers use.
Click Block to add a group that refuses a crawler outright. Compliance is voluntary - this stops cooperative bots, not scrapers.
| Token | What it is | Action |
|---|---|---|
| Googlebot | Google Search's main web crawler. Blocking it removes you from Google over time. | |
| Googlebot-Image | Crawls images for Google Images. Block it to keep photos out of image search. | |
| Bingbot | Microsoft Bing's crawler. Also feeds Bing-powered results in other products. | |
| DuckDuckBot | DuckDuckGo's own crawler for its instant answers and index. | |
| Baiduspider | Baidu's crawler, the dominant search engine in China. | |
| YandexBot | Yandex's crawler, the leading search engine in Russia. | |
| Slurp | Yahoo's legacy crawler. Still used for some Yahoo verticals. | |
| Applebot | Apple's crawler, powering Siri and Spotlight Suggestions. | |
| Twitterbot | Fetches link preview cards for X/Twitter. Blocking it kills your share previews. | |
| facebookexternalhit | Fetches Open Graph previews for Facebook, Messenger, and WhatsApp. | |
| GPTBotAI | OpenAI's crawler that collects public web content used to train its models. | |
| ClaudeBotAI | Anthropic's crawler for gathering training data for Claude. | |
| PerplexityBotAI | Perplexity's crawler, used to index pages its answer engine cites. | |
| CCBotAI | Common Crawl's bot. Its public corpus is a common source of AI training data. | |
| Google-ExtendedAI | Controls whether Google may use your content for Gemini and Vertex AI training. It does not affect Google Search ranking or indexing. | |
| BytespiderAI | ByteDance's crawler, associated with TikTok and its AI products. | |
| anthropic-aiAI | Legacy Anthropic token. Harmless to keep alongside ClaudeBot for older references. |

robots.txt 생성기는 도메인의 루트에 있는 일반 텍스트 파일을 빌드하고 크롤러에게 사이트의 어떤 부분을 가져올 수 있는지 알려줍니다. 형식은 단순해 보입니다. `allow:` 및 `disallow:` 규칙이 뒤에 오는 `allow:` 및 `disallow:` 룰이 있습니다. 하지만 세부 사항은 그룹 내에서만 적용되며, 경로는 슬래시로 시작되어야 하며, `*`와 `$`는 특정 방식으로 `*`와 `$`로 시작해야 하며, `user-agent: *`에서 전체 사이트를 Google에서 가져올 수 있습니다. 이 도구는 구문이 구성에 의해 올바르도록 구조화된 형식으로 파일을 빌드합니다.
그것은 생성을 넘어 세 가지 작업을 수행합니다. 파일을 린트하고 실제 풋건을 플래그 지정합니다. 그룹, 상대 경로, 맨발의 'Disallow:', 절대적으로 "모든 것을 허용", 절대적이지 않은 사이트 맵 URL, 알 수 없는 지시어 및 블랭킷 블록을 의미하는 맨손 '불허:`. 기존 robots.txt를 편집 가능한 그룹으로 다시 구문 분석하므로 현재 라이브를 붙여넣고 수정할 수 있습니다. 또한 실제 크롤러가 사용하는 것과 동일한 일치하는 가장 긴 일치 항목인 규칙에 대해 URL을 테스트합니다. 가장 구체적인 일치 패턴이 승리하고 비트가 연결을 허용하지 않도록 허용하므로 배포하기 전에 페이지에 도달할 수 있음을 증명할 수 있습니다.
분명히 말할 가치가 있는 한 가지는 robots.txt가 인덱싱이 아닌 크롤링을 제어하며 보안 경계가 아닙니다. Google은 가져온 적이 없는 URL을 색인을 생성할 수 있기 때문에 차단된 URL이 다른 페이지에 연결되어 있는 경우에도 검색 결과에 계속 표시될 수 있습니다. 인덱스에서 페이지를 유지하려면 'noindex' 메타 태그 또는 X-Robots-tag 헤더가 필요합니다. 크롤러는 페이지를 가져올 수 있는 경우에만 볼 수 있습니다. 그리고 파일은 공개되어 있습니다. /robots.txt에서 누구나 당신의 파일을 읽을 수 있으므로 거기에 비밀 디렉토리를 나열하면 광고합니다. 비공개에 대해 인증을 사용하십시오.
사전 설정을 선택하십시오. 모든 것을 허용하고, 모든 것을 차단하고, WordPress, Shopify, Next.js 또는 Block AI Crawlers와 같이 한 번의 클릭으로 양식을 채우거나 Live robots.txt를 붙여넣고 도구가 편집 가능한 그룹으로 다시 구문 분석합니다.
알려진 봇 목록에서 선택하거나 토큰을 입력하여 크롤러 또는 크롤러 세트당 그룹을 추가합니다. 각 그룹에 허용하지 않고 경로를 허용하고, 당신이 Honors를 목표로 하는 엔진이 크롤링 지연을 부여하십시오.
하나 이상의 절대 사이트맵 URL을 추가합니다. 검증자는 모든 키 입력에서 실행되며 오류, 경고 및 줄 번호가 있는 메모를 보고합니다. 배송 전에 빨간색을 수정합니다.
허용 또는 허용되지 않는 경로와 사용자 에이전트를 입력하여 정확히 어떤 규칙을 결정했는지 확인합니다. 예상대로 작동하면 파일을 복사하거나 robots.txt를 다운로드하여 도메인 루트에 업로드하십시오.
RFC 9309 우선 순위를 사용하여 모든 사용자 에이전트에 대해 경로를 테스트합니다. 가장 긴 일치 패턴이 승리하고 관계를 끊을 수 있습니다. 그리고 이를 결정한 정확한 규칙의 이름을 지정합니다.
그룹 외부 규칙, 상대 경로, 빈 허용 안 함, 절대 비절대 사이트맵 URL, 알 수 없는 지시어 및 블랭크 허용 안 함: / 사이트의 인덱싱을 거부합니다.
모든 것을 허용하고 모든 것을 차단하고 WordPress, Shopify, Next.js 및 블록 AI 크롤러가 각각 전체 양식을 합리적이고 정확한 시작점으로 채웁니다.
GPTBot, Claudebot, PerplexityBot, CCBot, Google-Extended, BytesPider, Anthropic-AI를 한 번의 클릭으로 차단하면서 검색 엔진을 완벽하게 환영합니다.
Free hreflang generator for multilingual and multi-region sites - build hreflang tags as HTML link tags, XML sitemap entries, or an HTTP header. Validates codes and x-default. 100% client-side.
Generate Apache .htaccess 301 and 302 redirect rules from a list of old-to-new URLs, plus force-HTTPS and www canonicalization directives, all in the browser.
Check one name across hundreds of TLDs at once. RDAP first, WHOIS fallback, and a verdict of available, taken or unknown — never a guess
Pull business listings for any search and place: name, address, category, rating, review count, phone and website, deduplicated and exportable as CSV.
0 comments
No comments yet. Be the first to share your thoughts!