Understanding the Challenge
AI labs like OpenAI claim their reasoning models outperform non-reasoning ones in specific areas. However, verifying these claims is hard due to the high costs of benchmarking. Evaluating reasoning models is expensive, making independent assessments challenging.
Key Insights
- OpenAI’s o1 reasoning model costs $2,767.05 to benchmark across seven tests.
- Anthropic’s Claude 3.7 Sonnet is cheaper at $1,485.35, while o3-mini costs $344.59.
- On average, reasoning models are more costly to evaluate than non-reasoning ones.
- Artificial Analysis has spent around $5,200 on reasoning models, nearly double the amount spent on over 80 non-reasoning models ($2,400).
The Bigger Picture
The rising costs of benchmarking reasoning models create barriers for independent verification. This situation affects transparency in AI research and development. As models become more complex, the expenses associated with testing increase, making it harder for smaller labs or academic institutions to participate. If results cannot be replicated due to high costs, the credibility of AI research may suffer. This could lead to a gap in understanding the true capabilities of these advanced models, limiting their potential impact in various fields.











