Understanding the New Bots
Meta has introduced two new bots, Meta-ExternalAgent and Meta-ExternalFetcher, aimed at collecting data to enhance its AI models. These bots are designed to make it difficult for website owners to block their content from being scraped. The Meta-ExternalAgent bot is intended for training AI models and indexing content, while the Meta-ExternalFetcher focuses on gathering web links to support AI assistant functions. These bots first emerged in July, raising concerns among website owners regarding the protection of their data.
Key Details
- The bots may bypass the traditional robots.txt rule, which website owners use to block automated scraping.
- Only 1.5% of top websites have managed to block the Meta-ExternalAgent bot, indicating its widespread acceptance.
- The Meta-ExternalFetcher is even less blocked, with less than 1% of top websites preventing its access.
- The dual functionality of the Meta-ExternalAgent bot complicates the ability for website owners to block data scraping while still allowing indexing for visibility.
Implications for Website Owners
The introduction of these bots highlights a growing tension between tech companies and website owners. As AI models require vast amounts of data, the traditional methods of protecting content are being challenged. Website owners may face difficulties in controlling how their data is used while still wanting their content to be discoverable. This issue raises important questions about data ownership and the balance between innovation in AI and the rights of content creators.











