Bots such as OpenAI’s GPTBot, the Applebot, CCBot, Google-Extended, and Bytespider analyze, store, or scrape your website’s data in order to provide data to train more advanced LLMs.
At Originality.ai, we care about responsible development (which includes ethical scraping) and the use (AI Detector) of Generative AI writing tools like ChatGPT.
This article will do a deep dive into the purpose of AI bots, what they do, how to block them and the interesting battle for the future of AI playing out on a rarely known file called the robots.txt.
AI bots come in multiple forms, including AI Assistants, AI Data Scrapers, and AI Search Crawlers all to power leading AI tools and AI search engines. Each of these AI bots extracts data from the web. Numerous webmasters find these practices unacceptable and want to keep their data or website information safe from being scraped.
As more of the internet is blocking AI bots, the number of words available to AI companies like OpenAI, Anthropic for developing their latest LLM (such as GPT-4o, Claude 3.5) will decline, resulting in slower future improvement of AI tools.
The most common way is to add the following text to your Robots.txt file:
User-agent: name-of-bot
Disallow: /
Example:
User-agent: GPTBot
Disallow: /
See the bottom of this article for a more in depth explanation of the options to block AI bots and a sample Robots.txt file.
OpenAI in particular, has been active in trying to secure data partnerships to continue to fuel its LLM training efforts.
Robots.txt is an internet protocol that provides search engine crawlers with information on which URLs the crawler can access on your website. It is primarily used to prevent overloading the domain with requests produced by crawlers.
It is important to know that Robots.txt is a request that bots should follow but does not “have” to be followed.
In terms of filtering and managing crawler traffic to your website, the purpose of robots.txt has a different application depending on the file type:
The leading purpose of robots.txt is to restrict crawler access to specific pages on your website. Suppose you have concerns that your website's server is overwhelmed by too many Google requests. In that case, you can prevent search engine crawlers from accessing specific pages on your website to reduce the utilization.
You can use the robots.txt file to manage the crawl traffic to images, videos, and audio files. This would prevent specific media from appearing in the SERP (Search Engine Results Page) on Google or other search engines.
If you think there's a particular utilization rate caused by unimportant style files, images, or scripts, you can use the robots.txt file to restrict the access of particular AI crawlers, scrapers, assistants or other bots.
Let’s review the different types of AI bots and crawlers deployed by companies on the web:
AI Assistants such as the ChatGPT-User owned by OpenAI at the Meta-ExternalFetcher deployed by Meta play a vital role in responding to user inquiries. The responses can either be in a text or voice format and use the collected web data to construct the most helpful answer possible to the user’s prompt.
AI web scraping is a procedure conducted by AI Data Scraper bots to harvest as much useful data as possible for LLM training. Companies such as Apple, ByteDance, Common Crawl, OpenAI and Anthropic use AI Data Scrapers to build a large dataset of the web for LLMs to train on.
Many companies deploy AI Search Crawlers to gather information about specific website pages, titles, keywords, images, and referenced inline links. While AI crawlers have the potential to send traffic to a website some website owners are choosing to still block it.
ChatGPT-User is a search assistant crawler dispatched by Open AI's ChatGPT as a result of user prompts. Most of its answers would typically include a summary of the website's content as well as a reference link.
The ChatGPT-User crawler's type is AI Assistant, as it is used to intelligently conduct tasks on behalf of the ChatGPT user.
The ChatGPT-User Search Assistant is expected to make one-off visits given the request of the user, rather than browsing the web automatically like other crawlers.
To block this crawler, you must include the following statement in the robots.txt of your website:
User-agent: ChatGPT-User
Disallow: /
The Meta-ExternalFetcher crawler is dispatched by all products of Meta AI to direct user prompts whenever an individual link is required.
The Meta-ExternalFetcher AI Assistant is fetched to intelligently perform tasks on behalf of the Meta AI user.
Similar to the ChatGPT-user crawler, the Meta-ExternalFetcher generally makes one-off visits based on the user's request, rather than automatically crawling the web.
You must include the following command in your website's robots.txt to prevent Meta-ExternalFetcher's access:
User-agent: Meta-ExternalFetcher
Disallow: /
The Amazonbot web crawler is used by Amazon to index and register search results which allows the Alexa AI Assistant to answer questions more accurately. Most of Alexa's answers generally contain a reference to the website.
Amazonbot is an AI Search Crawler that is used for indexing web content for Alexa's AI-powered search results.
The specific thing about search crawlers is that they don't adhere to a fixed visitation scheme for websites. The visitation frequency is defined by many factors and typically happens on-demand to a user query, including by the Amazonbot.
You can limit the Amazonbot's access to your website by typing down the following command lines in your website's robots.txt:
User-agent: Amazonbot
Disallow: /
The Applebot Search Crawler is used to register search results, allowing the Siri AI Assistant to answer user questions more effectively. Most of Siri's responses contain a reference to the websites crawled by Applebot.
Applebot is an AI Search Crawler that indexes web content to construct AI-powered search results.
The Applebot's behavior varies on multiple factors, such as search demand, crawled websites, and user queries. By default, search crawlers do not rely on fixed visitation to provide results.
While it is not advised to block search crawlers, you can use the following command in the website's robots.txt to prevent the Applebot's access:
User-agent: Applebot
Disallow: /
The OAI-SearchBot crawler is utilized to construct an index of websites that can be used as a result of OpenAI's SearchGPT product.
The OAI-SearchBot is an AI Search Crawler used for indexing web content to provide more accurate AI-powered search results for OpenAI's SearchGPT service.
The OAI-SearchBot's behavior can be defined by the frequency of web searches and user queries. Like any other search crawler, the OAI-SearchBot does not rely on fixed website visitation to provide results.
Include the following command in the robots.txt file of your website to actively prevent the OAI-SearchBot's access:
User-agent: OAI-SearchBot
Disallow: /
Perplexity uses the PerplexityBot web crawler to index search results for a more effective answer for their AI Assistant. The answers provided by the assistant normally surface inline references to a variety of web sources.
The PerplexityBot is an AI Search Crawler designed to index results for AI-powered search results by the Perplexity AI Assistant.
Like most search crawlers, the PerplexityBot does not depend on a fixed visitation schedule for the web sources it promotes. The frequency of visits can vary based on multiple factors, such as user queries.
You can restrict PerplexityBot's access to your website by including the following agent token rule in the robots.txt:
User-agent: PerplexityBot
Disallow: /
The YouBot is a search crawler deployed by You.ai to index search results for more accurate user answers by the You AI Assistant. The bot generally refers via inline sources to the referenced websites.
The YouBot Search Crawler indexes web content to generate more accurate AI-powered search results.
The YouBot crawler does not have a set visitation schedule and the frequency of visits often happens on-demand or in response to a user query.
You must paste the following command into your website's robots.txt file to prevent the YouBot crawler's access:
User-agent: YouBot
Disallow: /
The Applebot-Extended AI Data Scraper is used to train APple's lineup of LLM models that power the company's generative AI features. This Applebot scraper has a wide application in all aspects of Apple intelligence, Services, and Developer Tools.
The Applebot-Extended is an AI Data Scraper used for downloading web content and to train AI or LLM (Large Language Models)
While it remains unclear how exactly AI Data Scrapers choose which website to crawl, it is known that sources with a higher information density attract this scraper Applebot. It would make sense for an LLM to favor websites that regularly upload and update the on-page web information.
Include the following command in your website's robots.txt file to block the Applebot-Extended:
User-agent: Applebot-Extended
Disallow: /
Operated by ByteDance, Bytespider is an AI Data Scraper for the Chinese owner of TikTok. It's used to download LLM training data and supply relative data.
Bytespider is an AI Data Scraper used to train Large Language Models by downloading content from the web.
The Bytespider AI Data Scraper favors web sources with regularly updated and fact-rich information to use as a supply for LLMs.
Include the following use agent token rule in your website's robots.txt:
User-agent: Bytespider
Disallow: /
CCBot is owned by Common Crawl to construct an open-source repository through web crawl data available for anyone to access and use.
The CCBot is an AI Data Scraper purposed to download web content and conduct AI model training.
The CCBot crawls information-rich web sources to undergo more effective LLM training.
Include the following rule in robots.txt to restrict CCBot's access:
User-agent: CCBot
Disallow: /
The ClaudeBot AI Data Scraper is operated by Anthropic to supply Large Language Models like Claude with training data.
The ClaudeBot is an AI Data Scraper purposed for downloading web content and training AI models.
ClaudeBot chooses which websites to crawl based on the information density and the regularity of information updates.
The ClaudeBot's access can be revoked by including the following rule in the robots.txt:
User-agent: ClaudeBot
Disallow: /
The Diffbot is designed to structure, understand, aggregate, and even sell properly structured website data for AI model training and real-time monitoring.
The Diffbot is an AI Data Scraper designed to train AI models and download/structure web information.
The Diffbot's frequency of visitations is defined by the source's quality of information and updates regularity.
The Diffbot's crawl can be prevented by applying the following rule in the robots.txt:
User-agent: Diffbot
Disallow: /
The FacebookBot is deployed by Meta to enhance the AI speech recognition technology's efficiency and to train AI models.
FacebookBot is an AI Data Scraper crawler, used for registering web content and LLM training.
The FacebookBot does not have a fixed visitation schedule but perhaps recognizes and relies on sources with richer and well-updated information.
The FacebookBot's access can be revoked with the following rule:
User-agent: FacebookBot
Disallow: /
The Google-Extended crawler is used to supply training information for AI products owned by Google such as Gemini assistant and the Vertex AI generative APIs.
The Google-Extended crawler is an AI Data Scraper purposed to download information from the web and conduct AI training.
The Google-Extended bot’s visitation schedule is also flexible but it is much more directed than other crawlers due to Google's rich database of reliable information web sources.
The Google-Extended crawler can be blacklisted with the following rule:
User-agent: Google-Extended
Disallow: /
The GPTBot is developed by OpenAI to crawl web sources and download training data for the company's Large Language Models and products like ChatGPT.
The GPTBot is an AI Data Scraper designed to download and supply a wide range of data from the web.
Like other AI Data Scrapers, the GPTBot favors information-rich sources and websites to supply more relative information for AI training procedures.
You can block the GPTBot AI Data Scraper with the following robots.txt rule:
User-agent: GPTBot
Disallow: /
Meta-ExternalAgent एक क्रॉलर तकनीक है जिसे Meta ने वेब सामग्री को सीधे डाउनलोड और इंडेक्स करके कंपनी की AI तकनीकों को बेहतर बनाने के लिए विकसित किया है।
Meta-ExternalAgent वेब सामग्री को डाउनलोड और इंडेक्स करने के लिए AI डेटा स्क्रेपर तकनीक का उपयोग करता है, जिसका उद्देश्य AI प्रशिक्षण है।
कंपनी द्वारा विकसित अन्य क्रॉलर्स की तरह, Meta-ExternalAgent जानकारी-समृद्ध वेब स्रोतों को सटीक रूप से पहचानने के लिए एक लचीली क्रॉलिंग रणनीति का उपयोग करता है।
Meta द्वारा विकसित इस क्रॉलर को निम्नलिखित robots.txt नियम के माध्यम से प्रतिबंधित किया जा सकता है:
User-agent: Meta-ExternalAgent
Disallow: /
omgili क्रॉलर Webz.io के स्वामित्व में है, जिसे वेब क्रॉल डेटा की एक निर्मित लाइब्रेरी बनाए रखने के लिए डिज़ाइन किया गया है, जिसे बाद में AI प्रशिक्षण उद्देश्यों के लिए अन्य कंपनियों को बेचा जाता है।
omgili क्रॉलर एक AI डेटा स्क्रेपर है जो वेब से AI प्रशिक्षण जानकारी डाउनलोड करता है।
चूंकि क्रॉल की गई जानकारी बाद में Webz.io द्वारा बेची जाती है, omgili क्रॉलर संबंधित जानकारी वाली विश्वसनीय और अधिकृत वेबसाइटों का रिकॉर्ड रखता है।
omgili क्रॉलर की आपकी वेबसाइट तक पहुंच रोकने के लिए निम्नलिखित नियम का उपयोग करें:
User-agent: omgili
Disallow: /
अपुष्ट Anthropic-AI एजेंट का उपयोग संबंधित प्रशिक्षण डेटा डाउनलोड करने और उसे कंपनी के स्वामित्व वाले AI-संचालित उत्पादों, जैसे Claude, को उपलब्ध कराने के लिए किया जाता प्रतीत होता है।
कंपनी द्वारा खुलासा न किए जाने के कारण Anthropic-AI एजेंट का सटीक प्रकार अभी भी अज्ञात है।
Anthropic-AI एजेंट के बारे में संबंधित जानकारी की अनुपस्थिति के कारण, क्रॉलर का उपयोग कई उद्देश्यों के लिए किया जा सकता है, लेकिन अभी भी बताना कठिन है।
Anthropic-AI एजेंट को निम्नलिखित नियम से ब्लॉक किया जा सकता है:
User-agent: anthropic-ai
Disallow: /
Claude-Web Anthropic द्वारा संचालित एक और AI एजेंट है, जिसके उपयोग उद्देश्यों पर कोई आधिकारिक दस्तावेज़ नहीं है। Claude-Web से Anthropic के लिए संबंधित LLM प्रशिक्षण डेटा प्रदान करने की अपेक्षा की जाती है।
Claude-Web या तो AI डेटा स्क्रेपर होगा या Anthropic के Claude 3.5 Large Language Model के लिए एक मानक खोज क्रॉलर।
Anthropic Claude-Web की कार्यक्षमताओं के बारे में जानकारी रोक रहा है, लेकिन पूरी तरह खुलासा होने के बाद क्रॉलर का व्यवहार एजेंट के प्रकार के अनुरूप होगा।
आपकी वेबसाइट तक Claude-Web एजेंट की क्रॉलिंग पहुंच को निलंबित करने के लिए निम्नलिखित नियम का उपयोग किया जाता है:
User-agent: Claude-Web
Disallow: /
Cohere-AI एक अप्रलेखित एजेंट है जिसे Cohere ने अपने जनरेटिव AI टूल्स को संबंधित अध्ययन जानकारी प्रदान करने के लिए विकसित किया है। जब उपयोगकर्ता Cohere AI के माध्यम से संकेत देते हैं, तो यह वेब से जानकारी प्राप्त करता है।
चूंकि इस Cohere AI एजेंट के लिए कोई दस्तावेज़ उपलब्ध नहीं है, क्रॉलर का प्रकार अभी भी कई वेबसाइट मालिकों के लिए अज्ञात है।
संदेह है कि Cohere-AI एजेंट को विभिन्न व्यवहार पैटर्नों के माध्यम से बहुउद्देश्यीय बनाया जाएगा, ताकि Cohere उपयोगकर्ताओं को संबंधित जानकारी और इनलाइन स्रोत लिंक प्रदान किए जा सकें।
आप निम्नलिखित नियम के माध्यम से Cohere-AI Agent की पहुंच निलंबित कर सकते हैं:
User-agent: cohere-ai
Disallow: /
Ai2Bot का प्राथमिक कार्य “कुछ डोमेन्स” को क्रॉल करना और भाषा मॉडलों के प्रशिक्षण के लिए वेब सामग्री प्राप्त करना है।
Ai2 द्वारा रिपोर्ट किए अनुसार, Ai2Bot एक AI खोज क्रॉलर है, क्योंकि यह क्रॉल की गई वेबसाइट पर सामग्री, चित्रों और वीडियो का विश्लेषण करता है।
कंपनी द्वारा रिपोर्ट किए अनुसार Ai2Bot केवल विशिष्ट वेबसाइटों को क्रॉल करता है, लेकिन यह प्रतिदिन पंजीकृत डोमेन्स की अपनी सीमा बढ़ा सकता है।
Ai2Bot को निलंबित करने के लिए अपनी वेबसाइट के robots.txt में निम्नलिखित नियम शामिल करें:
User-agent: Ai2Bot
Disallow: /
Ai2 कंपनी Ai2Bot-Dolma बॉट की मालिक है और robots.txt नियमों का सम्मान करती है। प्राप्त सामग्री का उपयोग कंपनी के स्वामित्व वाले विभिन्न भाषा मॉडलों को प्रशिक्षित करने के लिए किया जाता है।
हालांकि बॉट को कोई विशिष्ट निर्धारित प्रकार नहीं दिया गया है, हमारा मानना है कि इसका व्यवहार एक मानक AI खोज क्रॉलर जैसा है।
Ai2Bot-Dolma भाषा मॉडलों के प्रशिक्षण के लिए आवश्यक वेब सामग्री खोजने हेतु केवल “कुछ डोमेन्स” को क्रॉल करता है।
Ai2Bot-Dolma की पहुंच प्रतिबंधित करने के लिए निम्नलिखित robots.txt पंक्ति का उपयोग करें:
User-agent: Ai2Bot-Dolma
Disallow: /
हालांकि इस क्रॉलर के बारे में अधिक जानकारी नहीं है, यह robots.txt का सम्मान करता है और मशीन लर्निंग प्रयोगों के लिए डेटा प्राप्त करने हेतु उपयोग किया जाता है।
बॉट का प्रकार अभी भी अज्ञात है, लेकिन इसके वेब व्यवहार के आधार पर हमारा मानना है कि यह या तो एक सामान्य क्रॉलर है या एक AI डेटा स्क्रेपर।
यह स्पष्ट नहीं है कि इस बॉट का ऑपरेटर कौन है, लेकिन डेटा का उपयोग बड़े भाषा मॉडल प्रशिक्षण, मशीन लर्निंग और डेटासेट निर्माण के लिए किया जाता है।
FriendlyCrawler की आपकी वेबसाइट तक पहुंच निलंबित करने का तरीका यह है:
User-agent: FriendlyCrawler
Disallow: /
GoogleOther क्रॉलर Google के स्वामित्व वाला एक बॉट है, लेकिन यह अभी भी स्पष्ट नहीं है कि यह AI-संबंधित है या बिल्कुल भी कृत्रिम रूप से बुद्धिमान है। वर्तमान में, शीर्ष प्रदर्शन करने वाली वेबसाइटों के केवल एक प्रतिशत ने GoogleOther बॉट को ब्लॉक किया है।
GoogleOther बॉट एक खोज इंजन क्रॉलर है जो अधिक सटीक खोज इंजन परिणामों या Google SERP के लिए वेब सामग्री को इंडेक्स करता है।
किसी भी अन्य खोज इंजन क्रॉलर की तरह, GoogleOther बॉट किसी विशेष विज़िट शेड्यूल का पालन नहीं करता और विज़िट आवृत्ति वेबसाइट गतिविधि और सामग्री गुणवत्ता पर निर्भर करती है।
अपनी वेबसाइट की robots.txt फ़ाइल में निम्नलिखित नियम जोड़ें:
User-agent: GoogleOther
Disallow: /
GoogleOther-Image GoogleOther का वह संस्करण है जिसे वेब पर चित्रों को क्रॉल, विश्लेषित और इंडेक्स करने के लिए प्रस्तावित किया गया है
GoogleOther की तरह, GoogleOther-Image बॉट एक सामान्य क्रॉलर है जिसका उपयोग विभिन्न उत्पाद टीमें वेबसाइट सामग्री को सार्वजनिक रूप से सुलभ बनाने के लिए करती हैं।
यह क्रॉलर किसी निश्चित विज़िट शेड्यूल से नहीं जुड़ा रहता, बल्कि वेब पर सबसे विश्वसनीय चित्र स्रोतों का विश्लेषण करता है और विशिष्ट जानकारी को इंडेक्स करता है।
आप इस Google क्रॉलर को निम्नलिखित कमांड से ब्लॉक कर सकते हैं:
User-agent: GoogleOther-Image
Disallow: /
मानक GoogleOther क्रॉलर और चित्र-पंजीकरण बॉट की तरह ही, GoogleOther-Video वेब पर वीडियो सामग्री को क्रॉल और विश्लेषित करता है।
यह बॉट एक सामान्य क्रॉलर है जो विभिन्न उत्पाद टीमों और व्यवसायों को उनकी वेबसाइट पर अपलोड किए गए वीडियो के साथ बेहतर पहुंच पाने में सहायता करता है।
GoogleOther के अन्य संस्करणों की तरह, इस क्रॉलर का भी एक लचीला विज़िट शेड्यूल है जो वेबसाइट वीडियो सामग्री की गतिविधि और गुणवत्ता से परिभाषित होता है।
इस Google बॉट को निम्नलिखित robots.txt पंक्ति से ब्लॉक किया जा सकता है:
User-agent: GoogleOther-Video
Disallow: /
ICC-Crawler एक एजेंट है जिसे निर्माता द्वारा अभी भी वर्गीकृत नहीं किया गया है। यह अभी भी अज्ञात है कि यह क्रॉलर कृत्रिम रूप से बुद्धिमान है या AI से इसका कोई संबंध है।
इसके स्रोत की तरह, ICC-Crawler बॉट का प्रकार भी अज्ञात है।
बॉट का व्यवहार क्रॉलर के प्रकार पर निर्भर करता है, विशेष रूप से इस पर कि वह डेटा स्क्रेपर, खोज इंजन क्रॉलर या आर्काइवर है।
ICC-Crawler को निम्नलिखित कमांड से ब्लॉक किया जा सकता है:
User-agent: ICC-Crawler
Disallow: /
Imagesift बॉट Hive के स्वामित्व में है, लेकिन वर्तमान में यह अज्ञात है कि क्रॉलर AI-संबंधित है या कृत्रिम रूप से बुद्धिमान है।
ImagesiftBot एक इंटेलिजेंस गैदरर है जो वेब पर उपयोगी इनसाइट्स खोजता है और परिणामों को डेटाबेस में पंजीकृत या इंडेक्स करता है।
इंटेलिजेंस गैदरर क्रॉलर्स का व्यवहार उनके ग्राहकों के लक्ष्यों पर निर्भर करता है। उदाहरण के लिए, कोई ग्राहक अपने ब्रांड को लोकप्रिय बनाने में रुचि रख सकता है, जिससे बॉट अन्य असंबंधित वेबसाइटों की तुलना में सोशल मीडिया को अधिक बार क्रॉल करता है।
इस क्रॉलर को निम्नलिखित तरीके से ब्लॉक किया जा सकता है:
User-agent: ImagesiftBot
Disallow: /
Huawei के स्वामित्व वाला PetalBot वर्तमान में लोकप्रिय इंडेक्स की गई वेबसाइटों में से 2% पर निलंबित है। यह अभी भी अज्ञात है कि यह क्रॉलर कृत्रिम रूप से बुद्धिमान है या किसी भी तरह AI से संबंधित है।
PetalBot एक खोज इंजन क्रॉलर है जो वेब सामग्री को इंडेक्स करता है और खोज इंजन परिणामों से डेटा प्राप्त करता है।
PetalBot का व्यवहार सामग्री की गुणवत्ता और पंजीकृत वेबसाइटों व डोमेन्स की गतिविधि से परिभाषित होता है। खोज इंजन क्रॉलर्स उच्च सामग्री गुणवत्ता और लगातार गतिविधि वाली वेबसाइटों को क्रॉल करने की प्रवृत्ति रखते हैं।
PetalBot की पहुंच निम्नलिखित तरीके से निलंबित की जा सकती है:
User-agent: PetalBot
Disallow: /
Scrapy Zyte के स्वामित्व में है और वर्तमान में वेब पर पंजीकृत डोमेन्स के 3% से अधिक पर ब्लॉक है।
Scrapy एक AI स्क्रेपर है, जो किसी वेबसाइट के robots.txt का सम्मान न करने के लिए कुख्यात क्रॉलर्स होते हैं। यह अंततः आवश्यक जानकारी का विश्लेषण करेगा, भले ही इसका मतलब ऐसी वेबसाइट तक पहुंचना हो जिसमें Scrapy के लिए disallow robots.txt नियम हो।
AI स्क्रेपर के विज़िट शेड्यूल की भविष्यवाणी करना लगभग असंभव है। इन क्रॉलर्स को अलग-अलग उद्देश्यों के साथ भेजा जाता है और यह बताना कठिन है कि वे किन वेबसाइटों को और कितनी बार क्रॉल करेंगे।
हालांकि robots.txt AI स्क्रेपर्स के विरुद्ध बहुत अधिक कुछ नहीं करता, Scrapy को ब्लॉक करने का तरीका यह है:
User-agent: Scrapy
Disallow: /
Timpibot Timpi के स्वामित्व में है और वर्तमान में लोकप्रिय इंडेक्स की गई वेबसाइटों में से 3% पर ब्लॉक है। Timpibot का एकमात्र उद्देश्य AI मॉडल प्रशिक्षण के लिए वेब डेटा प्राप्त करना है।
एक मानक स्क्रेपर के विपरीत, Timpibot एक AI डेटा स्क्रेपर है जो वेब सामग्री को क्रॉल और इंडेक्स करने के लिए पूरी तरह कृत्रिम बुद्धिमत्ता पर निर्भर करता है।
मानक स्क्रेपर्स की तरह, AI डेटा स्क्रेपर्स का विज़िट शेड्यूल भी अस्पष्ट है। ये क्रॉलर्स AI मॉडल प्रशिक्षण के लिए आवश्यक जानकारी के आधार पर उच्च जानकारी घनत्व और सामग्री मूल्य वाली वेबसाइटों को चुनने की प्रवृत्ति रखते हैं।
Timpibot की पहुंच निम्नलिखित robots.txt नियम से निलंबित की जा सकती है:
User-agent: Timpibot
Disallow: /
VelenPublicCrawler Hunter द्वारा संचालित है और वेब पर शीर्ष-इंडेक्स की गई वेबसाइटों में से 0% द्वारा ब्लॉक किया गया है।
क्रॉलर एक मानक इंटेलिजेंस गैदरर है, जिसका उद्देश्य वेब परिणामों से उपयोगी इनसाइट्स एकत्र करना है।
इंटेलिजेंस गैदरर्स अपने ग्राहकों के लक्ष्यों को पूरा करने की प्रवृत्ति रखते हैं और अधिकांश मामलों में, किस जानकारी को एकत्र करना है उसके लिए उनके पास विशिष्ट कार्य होते हैं।
VelenPublicWebCrawler को निम्नलिखित कमांड से ब्लॉक किया जा सकता है:
User-agent: VelenPublicWebCrawler
Disallow: /
Webzio-Extended Webz.io के स्वामित्व वाला एक और क्रॉलर है, जिसका उपयोग प्राप्त क्रॉल डेटा के रिपॉज़िटरी को बनाए रखने के लिए किया जाता है। जानकारी बाद में अन्य कंपनियों को बेची जाती है और आमतौर पर AI मॉडल प्रशिक्षण के लिए उपयोग की जाती है।
Webzio-Extended एक AI डेटा स्क्रेपर है जो AI मॉडलों को प्रशिक्षित करने के उद्देश्य से वेब सामग्री डाउनलोड करता है।
Webzio-Extended जैसे AI डेटा स्क्रेपर्स वेबसाइटों की निश्चित विज़िट पर कायम नहीं रहते। समृद्ध डेटा स्रोत स्क्रेपर्स को अधिक आकर्षित करते हैं और उन्हें अधिक बार क्रॉल करने का कारण बनते हैं।
Webzio-Extended की आपकी वेबसाइट तक पहुंच निम्नलिखित कमांड से निलंबित की जा सकती है:
User-agent: Webzio-Extended
Disallow: /
facebookexternalhit क्रॉलर Meta द्वारा भेजा गया एक बॉट है जो पंजीकृत लोकप्रिय वेबसाइटों के 6% से अधिक पर ब्लॉक है। यह अभी भी स्पष्ट नहीं है कि बॉट कृत्रिम रूप से बुद्धिमान है या AI से संबंधित है।
Facebookexternalhit एक फ़ेचर है, जो किसी एप्लिकेशन की ओर से वेब परिणामों को क्रॉल करता है।
फ़ेचर्स को आमतौर पर मांग पर वेबसाइटों पर जाने के लिए भेजा जाता है। उनका उपयोग किसी विशेष लिंक के मेटाडेटा, उदाहरण के लिए शीर्षक या थंबनेल छवि, को अनुरोध करने वाले उपयोगकर्ता के सामने प्रस्तुत करने के लिए किया जाता है।
इस Meta क्रॉलर को निम्नलिखित robots.txt नियम से ब्लॉक किया जा सकता है:
User-agent: facebookexternalhit
Disallow: /
यह ज्ञात है कि img2dataset कंपनी का एकमात्र वेब क्रॉलर है, जिसका उद्देश्य बड़ी संख्या में चित्र डाउनलोड करना और उन्हें बड़े भाषा मॉडलों को प्रशिक्षित करने के लिए डेटासेट में बदलना है।
img2dataset का प्रकार अभी भी अज्ञात है, लेकिन इसके व्यवहार और विज़िट शेड्यूल के आधार पर हमारा मानना है कि यह एक AI खोज इंजन क्रॉलर है।
Img2dataset अधिक समृद्ध चित्र डेटाबेस और विविध विषयों वाली वेबसाइटों को क्रॉल करने की प्रवृत्ति रखता है।
इस क्रॉलर को निम्नलिखित नियम से निलंबित किया जा सकता है:
User-agent: img2dataset
Disallow: /
आपकी वेबसाइट पर अवांछित बॉट ट्रैफ़िक को प्रतिबंधित करने की कई अनूठी संभावनाएँ हैं:
सबसे सामान्य तरीका अपनी Robots.txt फ़ाइल में निम्नलिखित पाठ जोड़ना है:
User-agent: name-of-bot
Disallow: /
उदाहरण:
User-agent: GPTBot
Disallow: /
आसानी से सुलभ robots.txt फ़ाइल बनाने में कई आसान चरण शामिल हैं:
इस चरण के दौरान, आपको वेबसाइट की रूट डायरेक्टरी तक पहुंचना होगा। आपकी वेबसाइट तक क्रॉलर पहुंच को प्रभावी रूप से प्रतिबंधित करने के लिए फ़ाइल को यहीं संग्रहीत किया जाना चाहिए। ध्यान रखें कि आपकी वेबसाइट में केवल एक robots.txt फ़ाइल हो सकती है।
आप फ़ाइल बनाने के लिए कोई भी एडिटर उपयोग कर सकते हैं, जैसे Windows का NotePad, TextEdit और vi। सुनिश्चित करें कि फ़ाइल UTF-8 एन्कोडिंग में सहेजी गई है और नियम लागू करने के लिए आगे बढ़ें। Google के क्रॉलर्स निम्नलिखित नियमों के सेट पर प्रतिक्रिया देते हैं: "user-agent," "disallow," "allow" और "sitemap." क्रॉलर्स को प्रबंधित करने के संदर्भ में प्रत्येक नियम का अपना अलग उद्देश्य होता है।
अगला चरण अपने ब्राउज़र में एक निजी ब्राउज़िंग विंडो खोलकर और फ़ाइल के स्थान पर नेविगेट करके नई बनाई गई robots.txt फ़ाइल को सार्वजनिक रूप से सुलभ बनाना है। आप साइट डोमेन टाइप करके, उदाहरण के लिए, "https://(site name)" और अंत में "/robots.txt" जोड़कर जांच सकते हैं कि robots.txt सार्वजनिक रूप से सुलभ है या नहीं।
जैसा कि हमने पहले स्थापित किया है, आपकी वेबसाइट के robots.txt में किसी विशेष बॉट के लिए प्रतिबंध नियम शामिल करने से न तो पेज डीइंडेक्स होगा और न ही SERP से हटेगा। आपके पेज तक पहुंचने का प्रयास करने पर, बॉट को स्वचालित रूप से एक त्रुटि संदेश दिया जाएगा और पेज की सामग्री तक पहुंच प्राप्त किए बिना उसे दूर रीडायरेक्ट कर दिया जाएगा।
पेज अभी भी सभी उपयोगकर्ताओं और AI बॉट्स के लिए खोजने योग्य रहेगा, लेकिन robots.txt फ़ाइल में सूचीबद्ध नेमटैग वाले बॉट्स को पहुंच नहीं मिलेगी।
यह उदाहरण फ़ाइल ब्लॉक करती है:
डेटा स्क्रेपर्स को ब्लॉक करना लेकिन क्रॉलर और खोज सहायक को नहीं, इसका उद्देश्य वेबसाइटों के डेटा को प्रशिक्षण के लिए उपयोग होने से प्रतिबंधित करना है, लेकिन AI खोज/सहायकों को वेबसाइट पर ट्रैफ़िक भेजने की अनुमति देना है।
फ़ायरवॉल की सहायता से AI बॉट्स की पहुंच निलंबित करने की बात आने पर कई संभावनाएँ होती हैं। आइए प्रत्येक अनूठी संभावना की समीक्षा करें:
यदि आप अपने डोमेन तक पहुंचने के लिए बॉट्स द्वारा उपयोग किए जाने वाले पतों से अच्छी तरह अवगत हैं, तो आप अपनी वेबसाइट के फ़ायरवॉल के माध्यम से IPs को ब्लैकलिस्ट कर सकते हैं। यह आपकी वेबसाइट पर अनपेक्षित ट्रैफ़िक को कम करने और संसाधनों को बचाने की एक सामान्य रूप से ज्ञात प्रथा है। हालांकि, बॉट्स कई IPs के माध्यम से चक्रित हो सकते हैं।
शायद आपकी वेबसाइट पर बॉट ट्रैफ़िक को 100% निलंबित करने के लिए सबसे व्यापक रूप से पसंदीदा तरीका ऐसा CAPTCHA सॉफ़्टवेयर लागू करना है जिसके लिए मानव सत्यापन आवश्यक हो। प्रत्येक नए विज़िटर को पहुंच प्राप्त करने के लिए पहेली के टुकड़ों का मिलान करने या चित्रों के सेट में वस्तुओं की पहचान करने जैसी एक सरल चुनौती पूरी करने के लिए कहा जाएगा। फ़ायरवॉल को CAPTCHAs ट्रिगर करने के लिए संशोधित किया जा सकता है।
Cloudflare या Amazon CloudFront जैसी CDN सेवाओं को AI बॉट्स के ट्रैफ़िक को कम करने के लिए आपकी वेबसाइट के फ़ायरवॉल के साथ एकीकृत किया जा सकता है। CAPTCHA चुनौतियों की तरह, केवल वास्तविक उपयोगकर्ताओं के स्वामित्व वाले वैध IP पते ही आपकी वेबसाइट तक पहुंचने की अनुमति पाएंगे।
चूंकि अनेक प्रकार के वेबमास्टर्स अपनी वेबसाइट पर अवांछित बॉट ट्रैफ़िक निलंबित करने के लिए robots.txt पर निर्भर करते हैं, फ़ाइल के भीतर प्रत्येक नियम किसी विशिष्ट बॉट के लिए सेट होता है। कंपनियों ने अपने बॉट्स के ट्रैफ़िक में गिरावट महसूस करनी शुरू की, तो कई ने अपने बॉट्स को नया नाम देने का निर्णय लिया जो कई वेबसाइटों की robots.txt फ़ाइल में नहीं था।
एक हालिया उदाहरण यह है कि Anthropic ने “ANTHROPIC-AI” और “CLAUDE-WEB” नामक अपने AI डेटा स्क्रेपर्स को “CLAUDEBOT” नामक एक नए बॉट में मिला दिया है। वेबसाइटों को इसके बारे में पता लगाने में निश्चित रूप से कुछ समय लगा और इस बीच, नए बॉट को इंटरनेट पर सभी वेबसाइटों तक अभूतपूर्व पहुंच मिल गई।
AI कंपनियों द्वारा बड़े भाषा मॉडलों को सिखाने के लिए सार्वजनिक जानकारी के बढ़ते और व्यापक उपयोग के साथ, वेब प्रोटोकॉल की अप्रभावशीलता स्पष्ट होती जा रही है। बड़े डेटा स्क्रैपिंग के जवाब में, जिसे C4, Dolma, और Refined Web जैसी वेब डेटा AI कंपनियों द्वारा किया गया, क्रॉलर पहुंच में 28%-45% गिरावट से अधिक हुई है।

C4 का पूरा 45% अब प्रतिबंधित हो चुका है, जिनमें से कई प्रतिबंध विविध हैं और robots.txt के माध्यम से सामान्य-उद्देश्य AI अवसंरचनाओं के नियमों का विस्तार कर रहे हैं। डेटा सहमति की मांग न केवल व्यावसायिक AI बल्कि सभी प्रकार के अकादमिक अनुसंधान और गैर-व्यावसायिक AI उपयोग के लिए अधिक से अधिक चुनौतीपूर्ण होती जा रही है।
LLMs को प्रशिक्षित करने के लिए जितना संभव हो उतना डेटा स्क्रैप करने की कोशिश कर रही AI कंपनियों और अपने डेटा/बैंडविड्थ को दुरुपयोग से बचाने के लिए काम कर रहे प्रकाशकों के टकराव के परिणामस्वरूप एक दिलचस्प संघर्ष उत्पन्न हुआ है, जो सभी वेबसाइटों पर robots.txt नामक एक कम-ज्ञात फ़ाइल में सामने आ रहा है।
इस गाइड का उद्देश्य हमारे लाइव डैशबोर्ड के साथ यह ट्रैक करना है कि यह ड्रामा कैसे सामने आ रहा है, यह देखते हुए कि शीर्ष 1000 वेबसाइटों में से कौन-सी AI बॉट्स को ब्लॉक कर रही हैं, उन्हें कैसे ब्लॉक करें और किन्हें आपको ब्लॉक करना चाहिए।
यदि हम कोई ऐसे बॉट्स छोड़ रहे हैं जिन्हें आप शामिल करना चाहेंगे, तो कृपया संपर्क करें।