Robots.txt का उपयोग करके AI बॉट्स को वेबसाइटों को क्रॉल करने से रोकें

लाइव डैशबोर्ड देखें, जो उन वेबसाइटों को दिखाता है जो GPTBot, CCBot, Google-extended और ByteSpider जैसे AI Bots को उनकी वेबसाइट पर मौजूद सामग्री को क्रॉल और स्क्रैप करने से रोक रही हैं। जानें कि कौन-से AI crawlers / scrapers क्या करते हैं और Robots.txt का उपयोग करके उन्हें कैसे ब्लॉक करें।

Bots such as OpenAI’s GPTBot, the Applebot, CCBot, Google-Extended, and Bytespider analyze, store, or scrape your website’s data in order to provide data to train more advanced LLMs. 

At Originality.ai, we care about responsible development (which includes ethical scraping) and the use (AI Detector) of Generative AI writing tools like ChatGPT.  

This article will do a deep dive into the purpose of AI bots, what they do, how to block them and the interesting battle for the future of AI playing out on a rarely known file called the robots.txt. 

What are AI Bots?

AI bots come in multiple forms, including AI Assistants, AI Data Scrapers, and AI Search Crawlers all to power leading AI tools and AI search engines. Each of these AI bots extracts data from the web. Numerous webmasters find these practices unacceptable and want to keep their data or website information safe from being scraped. 

Why Does it Matter if Websites Block AI Bots?

As more of the internet is blocking AI bots, the number of words available to AI companies like  OpenAI, Anthropic for developing their latest LLM (such as GPT-4o, Claude 3.5) will decline, resulting in slower future improvement of AI tools.

How To Block AI Bots (simple version):

The most common way is to add the following text to your Robots.txt file:

User-agent: name-of-bot

Disallow: /

Example:

User-agent: GPTBot

Disallow: /

See the bottom of this article for a more in depth explanation of the options to block AI bots and a sample Robots.txt file. 

OpenAI in particular, has been active in trying to secure data partnerships to continue to fuel its LLM training efforts. 

What is robots.txt and What Does it Do?

Robots.txt is an internet protocol that provides search engine crawlers with information on which URLs the crawler can access on your website. It is primarily used to prevent overloading the domain with requests produced by crawlers. 

It is important to know that Robots.txt is a request that bots should follow but does not “have” to be followed. 

In terms of filtering and managing crawler traffic to your website, the purpose of robots.txt has a different application depending on the file type:

Web Page Filtering

The leading purpose of robots.txt is to restrict crawler access to specific pages on your website. Suppose you have concerns that your website's server is overwhelmed by too many Google requests. In that case, you can prevent search engine crawlers from accessing specific pages on your website to reduce the utilization.

Media File Filtering

You can use the robots.txt file to manage the crawl traffic to images, videos, and audio files. This would prevent specific media from appearing in the SERP (Search Engine Results Page) on Google or other search engines.

Resource File 

If you think there's a particular utilization rate caused by unimportant style files, images, or scripts, you can use the robots.txt file to restrict the access of particular AI crawlers, scrapers, assistants or other bots.

Types of AI Bots

Let’s review the different types of AI bots and crawlers deployed by companies on the web:

AI Assistants

AI Assistants such as the ChatGPT-User owned by OpenAI at the Meta-ExternalFetcher deployed by Meta play a vital role in responding to user inquiries. The responses can either be in a text or voice format and use the collected web data to construct the most helpful answer possible to the user’s prompt.

AI Data Scrapers

AI web scraping is a procedure conducted by AI Data Scraper bots to harvest as much useful data as possible for LLM training. Companies such as Apple, ByteDance, Common Crawl, OpenAI and Anthropic use AI Data Scrapers to build a large dataset of the web for LLMs to train on.

AI Search Crawlers

Many companies deploy AI Search Crawlers to gather information about specific website pages, titles, keywords, images, and referenced inline links. While AI crawlers have the potential to send traffic to a website some website owners are choosing to still block it.

Description of Popular AI Bots

Deep Analysis of Popular AI Bots

ChatGPT-User - AI Search Assistant 

Overview

ChatGPT-User is a search assistant crawler dispatched by Open AI's ChatGPT as a result of user prompts. Most of its answers would typically include a summary of the website's content as well as a reference link.

Type

The ChatGPT-User crawler's type is AI Assistant, as it is used to intelligently conduct tasks on behalf of the ChatGPT user.

Crawler Behaviour

The ChatGPT-User Search Assistant is expected to make one-off visits given the request of the user, rather than browsing the web automatically like other crawlers.

How to Block ChatGPT-User Search Assistant?

To block this crawler, you must include the following statement in the robots.txt of your website:

User-agent: ChatGPT-User
Disallow: /

Meta-ExternalFetcher - AI Search Assistant 

Overview

The Meta-ExternalFetcher crawler is dispatched by all products of Meta AI to direct user prompts whenever an individual link is required.

Type

The Meta-ExternalFetcher AI Assistant is fetched to intelligently perform tasks on behalf of the Meta AI user.

Crawler Behaviour

Similar to the ChatGPT-user crawler, the Meta-ExternalFetcher generally makes one-off visits based on the user's request, rather than automatically crawling the web.

How to Block Meta-ExternalFetcher AI Assistant?

You must include the following command in your website's robots.txt to prevent Meta-ExternalFetcher's access:

User-agent: Meta-ExternalFetcher
Disallow: /

Amazonbot - AI Search Crawler 

Overview

The Amazonbot web crawler is used by Amazon to index and register search results which allows the Alexa AI Assistant to answer questions more accurately. Most of Alexa's answers generally contain a reference to the website.

Type

Amazonbot is an AI Search Crawler that is used for indexing web content for Alexa's AI-powered search results.

Crawler Behaviour

The specific thing about search crawlers is that they don't adhere to a fixed visitation scheme for websites. The visitation frequency is defined by many factors and typically happens on-demand to a user query, including by the Amazonbot.

How to Block Amazonbot Search Crawler?

You can limit the Amazonbot's access to your website by typing down the following command lines in your website's robots.txt:

User-agent: Amazonbot
Disallow: /

Applebot - AI Search Crawler 

Overview

The Applebot Search Crawler is used to register search results, allowing the Siri AI Assistant to answer user questions more effectively. Most of Siri's responses contain a reference to the websites crawled by Applebot.

Type

Applebot is an AI Search Crawler that indexes web content to construct AI-powered search results.

Crawler Behaviour

The Applebot's behavior varies on multiple factors, such as search demand, crawled websites, and user queries. By default, search crawlers do not rely on fixed visitation to provide results.

How to Block Applebot Search Crawler

While it is not advised to block search crawlers, you can use the following command in the website's robots.txt to prevent the Applebot's access:

User-agent: Applebot
Disallow: /

OAI-SearchBot - AI Search Crawler

Overview

The OAI-SearchBot crawler is utilized to construct an index of websites that can be used as a result of OpenAI's SearchGPT product.

Type

The OAI-SearchBot is an AI Search Crawler used for indexing web content to provide more accurate AI-powered search results for OpenAI's SearchGPT service.

Crawler Behaviour

The OAI-SearchBot's behavior can be defined by the frequency of web searches and user queries. Like any other search crawler, the OAI-SearchBot does not rely on fixed website visitation to provide results.

How to Block OAI-SearchBot Crawler?

Include the following command in the robots.txt file of your website to actively prevent the OAI-SearchBot's access:

User-agent: OAI-SearchBot
Disallow: /

PerplexityBot - AI Search Crawler

Overview

Perplexity uses the PerplexityBot web crawler to index search results for a more effective answer for their AI Assistant.  The answers provided by the assistant normally surface inline references to a variety of web sources.

Type

The PerplexityBot is an AI Search Crawler designed to index results for AI-powered search results by the Perplexity AI Assistant.

Crawler Behaviour

Like most search crawlers, the PerplexityBot does not depend on a fixed visitation schedule for the web sources it promotes. The frequency of visits can vary based on multiple factors, such as user queries.

How to Block PerplexityBot Search Crawler?

You can restrict PerplexityBot's access to your website by including the following agent token rule in the robots.txt:

User-agent: PerplexityBot
Disallow: /

YouBot - AI Search Crawler

Overview

The YouBot is a search crawler deployed by You.ai to index search results for more accurate user answers by the You AI Assistant. The bot generally refers via inline sources to the referenced websites.

Type

The YouBot Search Crawler indexes web content to generate more accurate AI-powered search results.

Crawler Behaviour

The YouBot crawler does not have a set visitation schedule and the frequency of visits often happens on-demand or in response to a user query.

How to Block YouBot Search Crawler?

You must paste the following command into your website's robots.txt file to prevent the YouBot crawler's access:

User-agent: YouBot
Disallow: /

Applebot-Extended AI Data Scraper

Overview

The Applebot-Extended AI Data Scraper is used to train APple's lineup of LLM models that power the company's generative AI features. This Applebot scraper has a wide application in all aspects of Apple intelligence, Services, and Developer Tools.

Type

The Applebot-Extended is an AI Data Scraper used for downloading web content and to train AI or LLM (Large Language Models)

Crawler Behaviour

While it remains unclear how exactly AI Data Scrapers choose which website to crawl, it is known that sources with a higher information density attract this scraper Applebot. It would make sense for an LLM to favor websites that regularly upload and update the on-page web information.

How to Block Applebot-Extended AI Data Scraper?

Include the following command in your website's robots.txt file to block the Applebot-Extended:

User-agent: Applebot-Extended
Disallow: /

Bytespider AI Data Scraper

Overview

Operated by ByteDance, Bytespider is an AI Data Scraper for the Chinese owner of TikTok. It's used to download LLM training data and supply relative data.

Type

Bytespider is an AI Data Scraper used to train Large Language Models by downloading content from the web.

Crawler Behaviour

The Bytespider AI Data Scraper favors web sources with regularly updated and fact-rich information to use as a supply for LLMs.

How to Block Bytespider AI Data Scraper?

Include the following use agent token rule in your website's robots.txt:

User-agent: Bytespider
Disallow: /

CCBot AI Data Scraper

Overview

CCBot is owned by Common Crawl to construct an open-source repository through web crawl data available for anyone to access and use.

Type

The CCBot is an AI Data Scraper purposed to download web content and conduct AI model training.

Crawler Behaviour

The CCBot crawls information-rich web sources to undergo more effective LLM training.

How to Block CCBot AI Data Scraper?

Include the following rule in robots.txt to restrict CCBot's access:

User-agent: CCBot
Disallow: /

ClaudeBot - AI Data Scraper

Overview

The ClaudeBot AI Data Scraper is operated by Anthropic to supply Large Language Models like Claude with training data.

Type

The ClaudeBot is an AI Data Scraper purposed for downloading web content and training AI models.

Crawler Behaviour

ClaudeBot chooses which websites to crawl based on the information density and the regularity of information updates.

How to Block ClaudeBot AI Data Scraper?

The ClaudeBot's access can be revoked by including the following rule in the robots.txt:

User-agent: ClaudeBot
Disallow: /

Diffbot - AI Data Scraper

Overview

The Diffbot is designed to structure, understand, aggregate, and even sell properly structured website data for AI model training and real-time monitoring.

Type

The Diffbot is an AI Data Scraper designed to train AI models and download/structure web information.

Crawler Behaviour

The Diffbot's frequency of visitations is defined by the source's quality of information and updates regularity.

How to Block Diffbot AI Data Scraper?

The Diffbot's crawl can be prevented by applying the following rule in the robots.txt:

User-agent: Diffbot
Disallow: /

FacebookBot - AI Data Scraper

Overview

The FacebookBot is deployed by Meta to enhance the AI speech recognition technology's efficiency and to train AI models.

Type

FacebookBot is an AI Data Scraper crawler, used for registering web content and LLM training.

Crawler Behaviour

The FacebookBot does not have a fixed visitation schedule but perhaps recognizes and relies on sources with richer and well-updated information.

How to Block FacebookBot AI Data Scraper?

The FacebookBot's access can be revoked with the following rule:

User-agent: FacebookBot
Disallow: /

Google-Extended - AI Data Scraper

Overview

The Google-Extended crawler is used to supply training information for AI products owned by Google such as Gemini assistant and the Vertex AI generative APIs.

Type

The Google-Extended crawler is an AI Data Scraper purposed to download information from the web and conduct AI training.

Crawler Behaviour

The Google-Extended bot’s visitation schedule is also flexible but it is much more directed than other crawlers due to Google's rich database of reliable information web sources.

How to Block Google-Extended AI Data Scraper?

The Google-Extended crawler can be blacklisted with the following rule:

User-agent: Google-Extended
Disallow: /

GPTBot - AI Data Scraper

Overview

The GPTBot is developed by OpenAI to crawl web sources and download training data for the company's Large Language Models and products like ChatGPT.

Type

The GPTBot is an AI Data Scraper designed to download and supply a wide range of data from the web.

Crawler Behaviour

Like other AI Data Scrapers, the GPTBot favors information-rich sources and websites to supply more relative information for AI training procedures.

How to Block GPTBot AI Data Scraper?

You can block the GPTBot AI Data Scraper with the following robots.txt rule:

User-agent: GPTBot
Disallow: /

Meta-ExternalAgent - AI डेटा स्क्रेपर

अवलोकन

Meta-ExternalAgent एक क्रॉलर तकनीक है जिसे Meta ने वेब सामग्री को सीधे डाउनलोड और इंडेक्स करके कंपनी की AI तकनीकों को बेहतर बनाने के लिए विकसित किया है।

प्रकार

Meta-ExternalAgent वेब सामग्री को डाउनलोड और इंडेक्स करने के लिए AI डेटा स्क्रेपर तकनीक का उपयोग करता है, जिसका उद्देश्य AI प्रशिक्षण है।

क्रॉलर व्यवहार

कंपनी द्वारा विकसित अन्य क्रॉलर्स की तरह, Meta-ExternalAgent जानकारी-समृद्ध वेब स्रोतों को सटीक रूप से पहचानने के लिए एक लचीली क्रॉलिंग रणनीति का उपयोग करता है।

Meta-ExternalAgent AI डेटा स्क्रेपर को कैसे ब्लॉक करें?

Meta द्वारा विकसित इस क्रॉलर को निम्नलिखित robots.txt नियम के माध्यम से प्रतिबंधित किया जा सकता है:

User-agent: Meta-ExternalAgent
Disallow: /

omgili - AI डेटा स्क्रेपर

अवलोकन

omgili क्रॉलर Webz.io के स्वामित्व में है, जिसे वेब क्रॉल डेटा की एक निर्मित लाइब्रेरी बनाए रखने के लिए डिज़ाइन किया गया है, जिसे बाद में AI प्रशिक्षण उद्देश्यों के लिए अन्य कंपनियों को बेचा जाता है।

प्रकार

omgili क्रॉलर एक AI डेटा स्क्रेपर है जो वेब से AI प्रशिक्षण जानकारी डाउनलोड करता है।

क्रॉलर व्यवहार

चूंकि क्रॉल की गई जानकारी बाद में Webz.io द्वारा बेची जाती है, omgili क्रॉलर संबंधित जानकारी वाली विश्वसनीय और अधिकृत वेबसाइटों का रिकॉर्ड रखता है।

omgili AI डेटा स्क्रेपर को कैसे ब्लॉक करें?

omgili क्रॉलर की आपकी वेबसाइट तक पहुंच रोकने के लिए निम्नलिखित नियम का उपयोग करें:

User-agent: omgili
Disallow: /

Anthropic-AI - अप्रलेखित AI एजेंट 

अवलोकन

अपुष्ट Anthropic-AI एजेंट का उपयोग संबंधित प्रशिक्षण डेटा डाउनलोड करने और उसे कंपनी के स्वामित्व वाले AI-संचालित उत्पादों, जैसे Claude, को उपलब्ध कराने के लिए किया जाता प्रतीत होता है।

प्रकार

कंपनी द्वारा खुलासा न किए जाने के कारण Anthropic-AI एजेंट का सटीक प्रकार अभी भी अज्ञात है।

क्रॉलर व्यवहार

Anthropic-AI एजेंट के बारे में संबंधित जानकारी की अनुपस्थिति के कारण, क्रॉलर का उपयोग कई उद्देश्यों के लिए किया जा सकता है, लेकिन अभी भी बताना कठिन है।

Anthropic-AI एजेंट को कैसे ब्लॉक करें?

Anthropic-AI एजेंट को निम्नलिखित नियम से ब्लॉक किया जा सकता है:

User-agent: anthropic-ai
Disallow: /

Claude-Web - अप्रलेखित AI एजेंट

अवलोकन

Claude-Web Anthropic द्वारा संचालित एक और AI एजेंट है, जिसके उपयोग उद्देश्यों पर कोई आधिकारिक दस्तावेज़ नहीं है। Claude-Web से Anthropic के लिए संबंधित LLM प्रशिक्षण डेटा प्रदान करने की अपेक्षा की जाती है।

प्रकार

Claude-Web या तो AI डेटा स्क्रेपर होगा या Anthropic के Claude 3.5 Large Language Model के लिए एक मानक खोज क्रॉलर।

क्रॉलर व्यवहार

Anthropic Claude-Web की कार्यक्षमताओं के बारे में जानकारी रोक रहा है, लेकिन पूरी तरह खुलासा होने के बाद क्रॉलर का व्यवहार एजेंट के प्रकार के अनुरूप होगा।

Claude-Web एजेंट को कैसे ब्लॉक करें?

आपकी वेबसाइट तक Claude-Web एजेंट की क्रॉलिंग पहुंच को निलंबित करने के लिए निम्नलिखित नियम का उपयोग किया जाता है:

User-agent: Claude-Web
Disallow: /

Cohere-AI एजेंट - अप्रलेखित

अवलोकन

Cohere-AI एक अप्रलेखित एजेंट है जिसे Cohere ने अपने जनरेटिव AI टूल्स को संबंधित अध्ययन जानकारी प्रदान करने के लिए विकसित किया है। जब उपयोगकर्ता Cohere AI के माध्यम से संकेत देते हैं, तो यह वेब से जानकारी प्राप्त करता है।

प्रकार

चूंकि इस Cohere AI एजेंट के लिए कोई दस्तावेज़ उपलब्ध नहीं है, क्रॉलर का प्रकार अभी भी कई वेबसाइट मालिकों के लिए अज्ञात है।

क्रॉलर व्यवहार

संदेह है कि Cohere-AI एजेंट को विभिन्न व्यवहार पैटर्नों के माध्यम से बहुउद्देश्यीय बनाया जाएगा, ताकि Cohere उपयोगकर्ताओं को संबंधित जानकारी और इनलाइन स्रोत लिंक प्रदान किए जा सकें।

Cohere-AI एजेंट को कैसे ब्लॉक करें?

आप निम्नलिखित नियम के माध्यम से Cohere-AI Agent की पहुंच निलंबित कर सकते हैं:

User-agent: cohere-ai
Disallow: /

Ai2Bot - AI खोज क्रॉलर

अवलोकन

Ai2Bot का प्राथमिक कार्य “कुछ डोमेन्स” को क्रॉल करना और भाषा मॉडलों के प्रशिक्षण के लिए वेब सामग्री प्राप्त करना है।

प्रकार

Ai2 द्वारा रिपोर्ट किए अनुसार, Ai2Bot एक AI खोज क्रॉलर है, क्योंकि यह क्रॉल की गई वेबसाइट पर सामग्री, चित्रों और वीडियो का विश्लेषण करता है।

क्रॉलर व्यवहार

कंपनी द्वारा रिपोर्ट किए अनुसार Ai2Bot केवल विशिष्ट वेबसाइटों को क्रॉल करता है, लेकिन यह प्रतिदिन पंजीकृत डोमेन्स की अपनी सीमा बढ़ा सकता है।

Ai2Bot AI खोज क्रॉलर को कैसे ब्लॉक करें?

Ai2Bot को निलंबित करने के लिए अपनी वेबसाइट के robots.txt में निम्नलिखित नियम शामिल करें:

User-agent: Ai2Bot

Disallow: /

Ai2Bot-Dolma - AI खोज क्रॉलर

अवलोकन

Ai2 कंपनी Ai2Bot-Dolma बॉट की मालिक है और robots.txt नियमों का सम्मान करती है। प्राप्त सामग्री का उपयोग कंपनी के स्वामित्व वाले विभिन्न भाषा मॉडलों को प्रशिक्षित करने के लिए किया जाता है।

प्रकार

हालांकि बॉट को कोई विशिष्ट निर्धारित प्रकार नहीं दिया गया है, हमारा मानना है कि इसका व्यवहार एक मानक AI खोज क्रॉलर जैसा है।

क्रॉलर व्यवहार

Ai2Bot-Dolma भाषा मॉडलों के प्रशिक्षण के लिए आवश्यक वेब सामग्री खोजने हेतु केवल “कुछ डोमेन्स” को क्रॉल करता है।

Ai2Bot-Dolma AI खोज क्रॉलर को कैसे ब्लॉक करें?

Ai2Bot-Dolma की पहुंच प्रतिबंधित करने के लिए निम्नलिखित robots.txt पंक्ति का उपयोग करें:

User-agent: Ai2Bot-Dolma

Disallow: /

FriendlyCrawler - अज्ञात

अवलोकन

हालांकि इस क्रॉलर के बारे में अधिक जानकारी नहीं है, यह robots.txt का सम्मान करता है और मशीन लर्निंग प्रयोगों के लिए डेटा प्राप्त करने हेतु उपयोग किया जाता है।

प्रकार

बॉट का प्रकार अभी भी अज्ञात है, लेकिन इसके वेब व्यवहार के आधार पर हमारा मानना है कि यह या तो एक सामान्य क्रॉलर है या एक AI डेटा स्क्रेपर।

क्रॉलर व्यवहार

यह स्पष्ट नहीं है कि इस बॉट का ऑपरेटर कौन है, लेकिन डेटा का उपयोग बड़े भाषा मॉडल प्रशिक्षण, मशीन लर्निंग और डेटासेट निर्माण के लिए किया जाता है।

FriendlyCrawler को कैसे ब्लॉक करें?

FriendlyCrawler की आपकी वेबसाइट तक पहुंच निलंबित करने का तरीका यह है:

User-agent: FriendlyCrawler

Disallow: /

GoogleOther - खोज इंजन क्रॉलर

अवलोकन

GoogleOther क्रॉलर Google के स्वामित्व वाला एक बॉट है, लेकिन यह अभी भी स्पष्ट नहीं है कि यह AI-संबंधित है या बिल्कुल भी कृत्रिम रूप से बुद्धिमान है। वर्तमान में, शीर्ष प्रदर्शन करने वाली वेबसाइटों के केवल एक प्रतिशत ने GoogleOther बॉट को ब्लॉक किया है।

प्रकार

GoogleOther बॉट एक खोज इंजन क्रॉलर है जो अधिक सटीक खोज इंजन परिणामों या Google SERP के लिए वेब सामग्री को इंडेक्स करता है।

क्रॉलर व्यवहार

किसी भी अन्य खोज इंजन क्रॉलर की तरह, GoogleOther बॉट किसी विशेष विज़िट शेड्यूल का पालन नहीं करता और विज़िट आवृत्ति वेबसाइट गतिविधि और सामग्री गुणवत्ता पर निर्भर करती है।

GoogleOther खोज इंजन क्रॉलर को कैसे ब्लॉक करें?

अपनी वेबसाइट की robots.txt फ़ाइल में निम्नलिखित नियम जोड़ें:

User-agent: GoogleOther

Disallow: /

GoogleOther-Image - सामान्य क्रॉलर

अवलोकन

GoogleOther-Image GoogleOther का वह संस्करण है जिसे वेब पर चित्रों को क्रॉल, विश्लेषित और इंडेक्स करने के लिए प्रस्तावित किया गया है

प्रकार

GoogleOther की तरह, GoogleOther-Image बॉट एक सामान्य क्रॉलर है जिसका उपयोग विभिन्न उत्पाद टीमें वेबसाइट सामग्री को सार्वजनिक रूप से सुलभ बनाने के लिए करती हैं। 

क्रॉलर व्यवहार

यह क्रॉलर किसी निश्चित विज़िट शेड्यूल से नहीं जुड़ा रहता, बल्कि वेब पर सबसे विश्वसनीय चित्र स्रोतों का विश्लेषण करता है और विशिष्ट जानकारी को इंडेक्स करता है।

GoogleOther-Image सामान्य क्रॉलर को कैसे ब्लॉक करें?

आप इस Google क्रॉलर को निम्नलिखित कमांड से ब्लॉक कर सकते हैं:

User-agent: GoogleOther-Image

Disallow: /

GoogleOther-Video - सामान्य क्रॉलर

अवलोकन

मानक GoogleOther क्रॉलर और चित्र-पंजीकरण बॉट की तरह ही, GoogleOther-Video वेब पर वीडियो सामग्री को क्रॉल और विश्लेषित करता है।

प्रकार

यह बॉट एक सामान्य क्रॉलर है जो विभिन्न उत्पाद टीमों और व्यवसायों को उनकी वेबसाइट पर अपलोड किए गए वीडियो के साथ बेहतर पहुंच पाने में सहायता करता है।

क्रॉलर व्यवहार

GoogleOther के अन्य संस्करणों की तरह, इस क्रॉलर का भी एक लचीला विज़िट शेड्यूल है जो वेबसाइट वीडियो सामग्री की गतिविधि और गुणवत्ता से परिभाषित होता है।

GoogleOther-Video सामान्य क्रॉलर को कैसे ब्लॉक करें?

इस Google बॉट को निम्नलिखित robots.txt पंक्ति से ब्लॉक किया जा सकता है:

User-agent: GoogleOther-Video

Disallow: /

ICC-Crawler - अज्ञात

अवलोकन

ICC-Crawler एक एजेंट है जिसे निर्माता द्वारा अभी भी वर्गीकृत नहीं किया गया है। यह अभी भी अज्ञात है कि यह क्रॉलर कृत्रिम रूप से बुद्धिमान है या AI से इसका कोई संबंध है।

प्रकार

इसके स्रोत की तरह, ICC-Crawler बॉट का प्रकार भी अज्ञात है।

क्रॉलर व्यवहार

बॉट का व्यवहार क्रॉलर के प्रकार पर निर्भर करता है, विशेष रूप से इस पर कि वह डेटा स्क्रेपर, खोज इंजन क्रॉलर या आर्काइवर है।

ICC-Crawler को कैसे ब्लॉक करें?

ICC-Crawler को निम्नलिखित कमांड से ब्लॉक किया जा सकता है:

User-agent: ICC-Crawler

Disallow: /

ImagesiftBot - इंटेलिजेंस गैदरर

अवलोकन

Imagesift बॉट Hive के स्वामित्व में है, लेकिन वर्तमान में यह अज्ञात है कि क्रॉलर AI-संबंधित है या कृत्रिम रूप से बुद्धिमान है। 

प्रकार

ImagesiftBot एक इंटेलिजेंस गैदरर है जो वेब पर उपयोगी इनसाइट्स खोजता है और परिणामों को डेटाबेस में पंजीकृत या इंडेक्स करता है।

क्रॉलर व्यवहार

इंटेलिजेंस गैदरर क्रॉलर्स का व्यवहार उनके ग्राहकों के लक्ष्यों पर निर्भर करता है। उदाहरण के लिए, कोई ग्राहक अपने ब्रांड को लोकप्रिय बनाने में रुचि रख सकता है, जिससे बॉट अन्य असंबंधित वेबसाइटों की तुलना में सोशल मीडिया को अधिक बार क्रॉल करता है।

ImagesiftBot इंटेलिजेंस गैदरर को कैसे ब्लॉक करें?

इस क्रॉलर को निम्नलिखित तरीके से ब्लॉक किया जा सकता है:

User-agent: ImagesiftBot

Disallow: /

PetalBot - खोज इंजन क्रॉलर

अवलोकन

Huawei के स्वामित्व वाला PetalBot वर्तमान में लोकप्रिय इंडेक्स की गई वेबसाइटों में से 2% पर निलंबित है। यह अभी भी अज्ञात है कि यह क्रॉलर कृत्रिम रूप से बुद्धिमान है या किसी भी तरह AI से संबंधित है।

प्रकार

PetalBot एक खोज इंजन क्रॉलर है जो वेब सामग्री को इंडेक्स करता है और खोज इंजन परिणामों से डेटा प्राप्त करता है।

क्रॉलर व्यवहार

PetalBot का व्यवहार सामग्री की गुणवत्ता और पंजीकृत वेबसाइटों व डोमेन्स की गतिविधि से परिभाषित होता है। खोज इंजन क्रॉलर्स उच्च सामग्री गुणवत्ता और लगातार गतिविधि वाली वेबसाइटों को क्रॉल करने की प्रवृत्ति रखते हैं।

PetalBot खोज इंजन क्रॉलर को कैसे ब्लॉक करें?

PetalBot की पहुंच निम्नलिखित तरीके से निलंबित की जा सकती है:

User-agent: PetalBot

Disallow: /

Scrapy - AI स्क्रेपर

अवलोकन

Scrapy Zyte के स्वामित्व में है और वर्तमान में वेब पर पंजीकृत डोमेन्स के 3% से अधिक पर ब्लॉक है।

प्रकार

Scrapy एक AI स्क्रेपर है, जो किसी वेबसाइट के robots.txt का सम्मान न करने के लिए कुख्यात क्रॉलर्स होते हैं। यह अंततः आवश्यक जानकारी का विश्लेषण करेगा, भले ही इसका मतलब ऐसी वेबसाइट तक पहुंचना हो जिसमें Scrapy के लिए disallow robots.txt नियम हो।

क्रॉलर व्यवहार

AI स्क्रेपर के विज़िट शेड्यूल की भविष्यवाणी करना लगभग असंभव है। इन क्रॉलर्स को अलग-अलग उद्देश्यों के साथ भेजा जाता है और यह बताना कठिन है कि वे किन वेबसाइटों को और कितनी बार क्रॉल करेंगे।

Scrapy AI स्क्रेपर को कैसे ब्लॉक करें?

हालांकि robots.txt AI स्क्रेपर्स के विरुद्ध बहुत अधिक कुछ नहीं करता, Scrapy को ब्लॉक करने का तरीका यह है:

User-agent: Scrapy

Disallow: /

Timpibot - AI डेटा स्क्रेपर

अवलोकन

Timpibot Timpi के स्वामित्व में है और वर्तमान में लोकप्रिय इंडेक्स की गई वेबसाइटों में से 3% पर ब्लॉक है। Timpibot का एकमात्र उद्देश्य AI मॉडल प्रशिक्षण के लिए वेब डेटा प्राप्त करना है।

प्रकार

एक मानक स्क्रेपर के विपरीत, Timpibot एक AI डेटा स्क्रेपर है जो वेब सामग्री को क्रॉल और इंडेक्स करने के लिए पूरी तरह कृत्रिम बुद्धिमत्ता पर निर्भर करता है। 

क्रॉलर व्यवहार

मानक स्क्रेपर्स की तरह, AI डेटा स्क्रेपर्स का विज़िट शेड्यूल भी अस्पष्ट है। ये क्रॉलर्स AI मॉडल प्रशिक्षण के लिए आवश्यक जानकारी के आधार पर उच्च जानकारी घनत्व और सामग्री मूल्य वाली वेबसाइटों को चुनने की प्रवृत्ति रखते हैं।

Timpibot AI डेटा स्क्रेपर को कैसे ब्लॉक करें?

Timpibot की पहुंच निम्नलिखित robots.txt नियम से निलंबित की जा सकती है:

User-agent: Timpibot

Disallow: /

VelenPublicWebCrawler - इंटेलिजेंस गैदरर

अवलोकन

VelenPublicCrawler Hunter द्वारा संचालित है और वेब पर शीर्ष-इंडेक्स की गई वेबसाइटों में से 0% द्वारा ब्लॉक किया गया है।

प्रकार

क्रॉलर एक मानक इंटेलिजेंस गैदरर है, जिसका उद्देश्य वेब परिणामों से उपयोगी इनसाइट्स एकत्र करना है।

क्रॉलर व्यवहार

इंटेलिजेंस गैदरर्स अपने ग्राहकों के लक्ष्यों को पूरा करने की प्रवृत्ति रखते हैं और अधिकांश मामलों में, किस जानकारी को एकत्र करना है उसके लिए उनके पास विशिष्ट कार्य होते हैं।

VelenPublicWebCrawler इंटेलिजेंस गैदरर को कैसे ब्लॉक करें?

VelenPublicWebCrawler को निम्नलिखित कमांड से ब्लॉक किया जा सकता है:

User-agent: VelenPublicWebCrawler

Disallow: /

Webzio-Extended - AI डेटा स्क्रेपर

अवलोकन

Webzio-Extended Webz.io के स्वामित्व वाला एक और क्रॉलर है, जिसका उपयोग प्राप्त क्रॉल डेटा के रिपॉज़िटरी को बनाए रखने के लिए किया जाता है। जानकारी बाद में अन्य कंपनियों को बेची जाती है और आमतौर पर AI मॉडल प्रशिक्षण के लिए उपयोग की जाती है।

प्रकार

Webzio-Extended एक AI डेटा स्क्रेपर है जो AI मॉडलों को प्रशिक्षित करने के उद्देश्य से वेब सामग्री डाउनलोड करता है।

क्रॉलर व्यवहार

Webzio-Extended जैसे AI डेटा स्क्रेपर्स वेबसाइटों की निश्चित विज़िट पर कायम नहीं रहते। समृद्ध डेटा स्रोत स्क्रेपर्स को अधिक आकर्षित करते हैं और उन्हें अधिक बार क्रॉल करने का कारण बनते हैं।

Webzio-Extended AI डेटा स्क्रेपर को कैसे ब्लॉक करें?

Webzio-Extended की आपकी वेबसाइट तक पहुंच निम्नलिखित कमांड से निलंबित की जा सकती है:

User-agent: Webzio-Extended

Disallow: /

Facebookexternalhit - फ़ेचर

अवलोकन

facebookexternalhit क्रॉलर Meta द्वारा भेजा गया एक बॉट है जो पंजीकृत लोकप्रिय वेबसाइटों के 6% से अधिक पर ब्लॉक है। यह अभी भी स्पष्ट नहीं है कि बॉट कृत्रिम रूप से बुद्धिमान है या AI से संबंधित है।

प्रकार

Facebookexternalhit एक फ़ेचर है, जो किसी एप्लिकेशन की ओर से वेब परिणामों को क्रॉल करता है। 

क्रॉलर व्यवहार

फ़ेचर्स को आमतौर पर मांग पर वेबसाइटों पर जाने के लिए भेजा जाता है। उनका उपयोग किसी विशेष लिंक के मेटाडेटा, उदाहरण के लिए शीर्षक या थंबनेल छवि, को अनुरोध करने वाले उपयोगकर्ता के सामने प्रस्तुत करने के लिए किया जाता है।

Facebookexternalhit फ़ेचर को कैसे ब्लॉक करें?

इस Meta क्रॉलर को निम्नलिखित robots.txt नियम से ब्लॉक किया जा सकता है:

User-agent: facebookexternalhit 

Disallow: /

Img2dataset - अज्ञात

अवलोकन

यह ज्ञात है कि img2dataset कंपनी का एकमात्र वेब क्रॉलर है, जिसका उद्देश्य बड़ी संख्या में चित्र डाउनलोड करना और उन्हें बड़े भाषा मॉडलों को प्रशिक्षित करने के लिए डेटासेट में बदलना है।

प्रकार

img2dataset का प्रकार अभी भी अज्ञात है, लेकिन इसके व्यवहार और विज़िट शेड्यूल के आधार पर हमारा मानना है कि यह एक AI खोज इंजन क्रॉलर है।

क्रॉलर व्यवहार

Img2dataset अधिक समृद्ध चित्र डेटाबेस और विविध विषयों वाली वेबसाइटों को क्रॉल करने की प्रवृत्ति रखता है।

img2dataset क्रॉलर को कैसे ब्लॉक करें?

इस क्रॉलर को निम्नलिखित नियम से निलंबित किया जा सकता है:

User-agent: img2dataset

Disallow: /

AI बॉट्स को कैसे ब्लॉक करें (उन्नत संस्करण):

आपकी वेबसाइट पर अवांछित बॉट ट्रैफ़िक को प्रतिबंधित करने की कई अनूठी संभावनाएँ हैं:

सबसे सामान्य तरीका अपनी Robots.txt फ़ाइल में निम्नलिखित पाठ जोड़ना है:

User-agent: name-of-bot
Disallow: /

उदाहरण:

User-agent: GPTBot
Disallow: /

Robots.txt का उपयोग करें

आसानी से सुलभ robots.txt फ़ाइल बनाने में कई आसान चरण शामिल हैं:

1. "robots.txt" नाम की फ़ाइल बनाएँ

इस चरण के दौरान, आपको वेबसाइट की रूट डायरेक्टरी तक पहुंचना होगा। आपकी वेबसाइट तक क्रॉलर पहुंच को प्रभावी रूप से प्रतिबंधित करने के लिए फ़ाइल को यहीं संग्रहीत किया जाना चाहिए। ध्यान रखें कि आपकी वेबसाइट में केवल एक robots.txt फ़ाइल हो सकती है।

2. robots.txt नियम लिखें

आप फ़ाइल बनाने के लिए कोई भी एडिटर उपयोग कर सकते हैं, जैसे Windows का NotePad, TextEdit और vi। सुनिश्चित करें कि फ़ाइल UTF-8 एन्कोडिंग में सहेजी गई है और नियम लागू करने के लिए आगे बढ़ें।  Google के क्रॉलर्स निम्नलिखित नियमों के सेट पर प्रतिक्रिया देते हैं: "user-agent," "disallow," "allow" और "sitemap." क्रॉलर्स को प्रबंधित करने के संदर्भ में प्रत्येक नियम का अपना अलग उद्देश्य होता है।

3. robots.txt फ़ाइल अपलोड करें

अगला चरण अपने ब्राउज़र में एक निजी ब्राउज़िंग विंडो खोलकर और फ़ाइल के स्थान पर नेविगेट करके नई बनाई गई robots.txt फ़ाइल को सार्वजनिक रूप से सुलभ बनाना है। आप साइट डोमेन टाइप करके, उदाहरण के लिए, "https://(site name)" और अंत में "/robots.txt" जोड़कर जांच सकते हैं कि robots.txt सार्वजनिक रूप से सुलभ है या नहीं।

AI डेटा स्क्रेपर बॉट्स को ब्लॉक करने वाला नमूना Robots.txt

जैसा कि हमने पहले स्थापित किया है, आपकी वेबसाइट के robots.txt में किसी विशेष बॉट के लिए प्रतिबंध नियम शामिल करने से न तो पेज डीइंडेक्स होगा और न ही SERP से हटेगा। आपके पेज तक पहुंचने का प्रयास करने पर, बॉट को स्वचालित रूप से एक त्रुटि संदेश दिया जाएगा और पेज की सामग्री तक पहुंच प्राप्त किए बिना उसे दूर रीडायरेक्ट कर दिया जाएगा।

पेज अभी भी सभी उपयोगकर्ताओं और AI बॉट्स के लिए खोजने योग्य रहेगा, लेकिन robots.txt फ़ाइल में सूचीबद्ध नेमटैग वाले बॉट्स को पहुंच नहीं मिलेगी।

यह उदाहरण फ़ाइल ब्लॉक करती है:

  • नहीं - AI खोज सहायक
  • नहीं - AI खोज क्रॉलर
  • सभी - AI डेटा स्क्रेपर

डेटा स्क्रेपर्स को ब्लॉक करना लेकिन क्रॉलर और खोज सहायक को नहीं, इसका उद्देश्य वेबसाइटों के डेटा को प्रशिक्षण के लिए उपयोग होने से प्रतिबंधित करना है, लेकिन AI खोज/सहायकों को वेबसाइट पर ट्रैफ़िक भेजने की अनुमति देना है।

फ़ायरवॉल

फ़ायरवॉल की सहायता से AI बॉट्स की पहुंच निलंबित करने की बात आने पर कई संभावनाएँ होती हैं। आइए प्रत्येक अनूठी संभावना की समीक्षा करें:

  • IP ब्लॉकिंग सेट अप करें

यदि आप अपने डोमेन तक पहुंचने के लिए बॉट्स द्वारा उपयोग किए जाने वाले पतों से अच्छी तरह अवगत हैं, तो आप अपनी वेबसाइट के फ़ायरवॉल के माध्यम से IPs को ब्लैकलिस्ट कर सकते हैं। यह आपकी वेबसाइट पर अनपेक्षित ट्रैफ़िक को कम करने और संसाधनों को बचाने की एक सामान्य रूप से ज्ञात प्रथा है। हालांकि, बॉट्स कई IPs के माध्यम से चक्रित हो सकते हैं।

  • CAPTCHA सॉफ़्टवेयर का उपयोग करें

शायद आपकी वेबसाइट पर बॉट ट्रैफ़िक को 100% निलंबित करने के लिए सबसे व्यापक रूप से पसंदीदा तरीका ऐसा CAPTCHA सॉफ़्टवेयर लागू करना है जिसके लिए मानव सत्यापन आवश्यक हो। प्रत्येक नए विज़िटर को पहुंच प्राप्त करने के लिए पहेली के टुकड़ों का मिलान करने या चित्रों के सेट में वस्तुओं की पहचान करने जैसी एक सरल चुनौती पूरी करने के लिए कहा जाएगा। फ़ायरवॉल को CAPTCHAs ट्रिगर करने के लिए संशोधित किया जा सकता है

CDN (Content Delivery Network) का उपयोग करें

Cloudflare या Amazon CloudFront जैसी CDN सेवाओं को AI बॉट्स के ट्रैफ़िक को कम करने के लिए आपकी वेबसाइट के फ़ायरवॉल के साथ एकीकृत किया जा सकता है। CAPTCHA चुनौतियों की तरह, केवल वास्तविक उपयोगकर्ताओं के स्वामित्व वाले वैध IP पते ही आपकी वेबसाइट तक पहुंचने की अनुमति पाएंगे।

बॉट नाम बदलने की उभरती चुनौतियाँ

चूंकि अनेक प्रकार के वेबमास्टर्स अपनी वेबसाइट पर अवांछित बॉट ट्रैफ़िक निलंबित करने के लिए robots.txt पर निर्भर करते हैं, फ़ाइल के भीतर प्रत्येक नियम किसी विशिष्ट बॉट के लिए सेट होता है। कंपनियों ने अपने बॉट्स के ट्रैफ़िक में गिरावट महसूस करनी शुरू की, तो कई ने अपने बॉट्स को नया नाम देने का निर्णय लिया जो कई वेबसाइटों की robots.txt फ़ाइल में नहीं था।

एक हालिया उदाहरण यह है कि Anthropic ने “ANTHROPIC-AI” और “CLAUDE-WEB” नामक अपने AI डेटा स्क्रेपर्स को “CLAUDEBOT” नामक एक नए बॉट में मिला दिया है। वेबसाइटों को इसके बारे में पता लगाने में निश्चित रूप से कुछ समय लगा और इस बीच, नए बॉट को इंटरनेट पर सभी वेबसाइटों तक अभूतपूर्व पहुंच मिल गई।

AI डेटा कॉमन्स में गिरावट और परिणाम

AI कंपनियों द्वारा बड़े भाषा मॉडलों को सिखाने के लिए सार्वजनिक जानकारी के बढ़ते और व्यापक उपयोग के साथ, वेब प्रोटोकॉल की अप्रभावशीलता स्पष्ट होती जा रही है। बड़े डेटा स्क्रैपिंग के जवाब में, जिसे C4, Dolma, और Refined Web जैसी वेब डेटा AI कंपनियों द्वारा किया गया, क्रॉलर पहुंच में 28%-45% गिरावट से अधिक हुई है।

Decline of AI Data Commons and Consequences
https://www.dataprovenance.org/Consent_in_Crisis.pdf

C4 का पूरा 45% अब प्रतिबंधित हो चुका है, जिनमें से कई प्रतिबंध विविध हैं और robots.txt के माध्यम से सामान्य-उद्देश्य AI अवसंरचनाओं के नियमों का विस्तार कर रहे हैं। डेटा सहमति की मांग न केवल व्यावसायिक AI बल्कि सभी प्रकार के अकादमिक अनुसंधान और गैर-व्यावसायिक AI उपयोग के लिए अधिक से अधिक चुनौतीपूर्ण होती जा रही है।

निष्कर्ष

LLMs को प्रशिक्षित करने के लिए जितना संभव हो उतना डेटा स्क्रैप करने की कोशिश कर रही AI कंपनियों और अपने डेटा/बैंडविड्थ को दुरुपयोग से बचाने के लिए काम कर रहे प्रकाशकों के टकराव के परिणामस्वरूप एक दिलचस्प संघर्ष उत्पन्न हुआ है, जो सभी वेबसाइटों पर robots.txt नामक एक कम-ज्ञात फ़ाइल में सामने आ रहा है। 

इस गाइड का उद्देश्य हमारे लाइव डैशबोर्ड के साथ यह ट्रैक करना है कि यह ड्रामा कैसे सामने आ रहा है, यह देखते हुए कि शीर्ष 1000 वेबसाइटों में से कौन-सी AI बॉट्स को ब्लॉक कर रही हैं, उन्हें कैसे ब्लॉक करें और किन्हें आपको ब्लॉक करना चाहिए। 

यदि हम कोई ऐसे बॉट्स छोड़ रहे हैं जिन्हें आप शामिल करना चाहेंगे, तो कृपया संपर्क करें।

Al कंटेंट डिटेक्टर & प्लेजरिज़्म चेकर मार्केटर्स और लेखकों के लिए

हमारे अग्रणी टूल्स का उपयोग करें ताकि आप ईमानदारी के साथ प्रकाशित कर सकें!