The Danish media outlets’ demand to remove their articles from Common Crawl’s data sets has sparked a heated debate about copyrighted materials and artificial intelligence. The nonprofit web archive, Common Crawl, plans to comply with the request, citing its lack of resources to fight media companies in court. This move is seen as a significant blow to AI development, as Common Crawl’s data has been instrumental in training many text-based generative AI tools. The Danish Rights Alliance, representing copyright holders in Denmark, led the campaign, inspired by The New York Times’ similar request last year. The alliance argues that Common Crawl’s corpus poses a threat to media companies negotiating with AI giants. However, Common Crawl’s executive director, Rich Skrenta, views this push to remove archival materials as an existential threat to the open web. Amid growing outrage, it’s clear that the battle over AI’s data sources is only just beginning.

Source.

TOP STORIES

Democrats Urged to Prioritize AI Safety and Economic Impact
Obama stresses Democrats must prioritize AI safety and economic strategy …
Pacing AI Development - A Call for Caution from Industry Leaders
Amodei’s call for caution in AI development highlights the need for safety and alignment …
Big Tech's Trust Crisis Deepens with Anthropic Lawsuit
Sony Music and Warner Music have sued Anthropic, accusing it of copyright infringement in AI training …
Nvidia's AI Future - Jensen Huang's Vision for Record Growth
Huang believes Nvidia’s position in AI will lead to another year of record growth …
China's AI Companies Target US Models with Distillation Attacks
Anthropic’s report reveals a surge in distillation attacks by Chinese AI firms on U.S. models …
Cybersecurity Concerns Rise as AI Agents Break Boundaries
AI agents’ autonomy poses significant risks, as demonstrated by a recent breach …

latest stories