Anthropic has introduced a program aimed at funding the creation of new benchmarks to evaluate AI models, including generative models like Claude. This initiative will allocate resources to third-party organizations capable of developing effective measures for advanced AI capabilities. Anthropic emphasizes the growing need for high-quality, safety-focused evaluations that can keep pace with the rapid advancements in AI. The company is particularly interested in benchmarks that assess AI’s ability to execute tasks with significant societal and security implications, such as cyberattacks, weapons enhancement, and misinformation spread. Additionally, Anthropic aims to support research into benchmarks that examine AI’s potential in scientific research, multilingual communication, and bias mitigation, among other areas. The program will feature a range of funding options and involve collaboration with Anthropic’s domain experts. While the initiative has noble goals, its success may depend on the level of funding and manpower committed. Critics, however, may question Anthropic’s definitions of “safe” and “risky” AI and the company’s commercial motives. Despite these concerns, Anthropic aspires for its program to set a new industry standard for comprehensive AI evaluation.

Source.

TOP STORIES

Democrats Urged to Prioritize AI Safety and Economic Impact
Obama stresses Democrats must prioritize AI safety and economic strategy …
Pacing AI Development - A Call for Caution from Industry Leaders
Amodei’s call for caution in AI development highlights the need for safety and alignment …
Big Tech's Trust Crisis Deepens with Anthropic Lawsuit
Sony Music and Warner Music have sued Anthropic, accusing it of copyright infringement in AI training …
Nvidia's AI Future - Jensen Huang's Vision for Record Growth
Huang believes Nvidia’s position in AI will lead to another year of record growth …
China's AI Companies Target US Models with Distillation Attacks
Anthropic’s report reveals a surge in distillation attacks by Chinese AI firms on U.S. models …
Cybersecurity Concerns Rise as AI Agents Break Boundaries
AI agents’ autonomy poses significant risks, as demonstrated by a recent breach …

latest stories