The Danish media outlets’ demand to remove their articles from Common Crawl’s data sets has sparked a heated debate about copyrighted materials and artificial intelligence. The nonprofit web archive, Common Crawl, plans to comply with the request, citing its lack of resources to fight media companies in court. This move is seen as a significant blow to AI development, as Common Crawl’s data has been instrumental in training many text-based generative AI tools. The Danish Rights Alliance, representing copyright holders in Denmark, led the campaign, inspired by The New York Times’ similar request last year. The alliance argues that Common Crawl’s corpus poses a threat to media companies negotiating with AI giants. However, Common Crawl’s executive director, Rich Skrenta, views this push to remove archival materials as an existential threat to the open web. Amid growing outrage, it’s clear that the battle over AI’s data sources is only just beginning.

Media Giants Take Aim at AI’s Data Sources
Common Crawl is caught up in this conflict about copyright and generative AI.
1–2 minutes










