🤖AI Index v4.1.1 Updates Methodology and Scores
New models join the leaderboard, but costs are up
TL;DR
The Artificial Analysis Intelligence Index v4.1.1 introduces updates to its methodology and scores new language models like Claude Opus 5 and DeepSeek V4 Flash. The cost per task metric now shows slight absolute increases but maintains relative positioning.
Artificial Analysis Intelligence Index v4.1.1 has rolled out with a fresh batch of evaluations, including the latest models from Ling, Muse Spark, Qwen3.8 Max, and Agnes AI's Agnes 2.5 Pro Alpha. The update also tweaks the Cost per Task metric to reflect more accurate pricing data, though it doesn't significantly alter model rankings. If you're in the market for a cost-effective language model with strong performance, Claude Opus 5 stands out as a new leader in agentic knowledge work, offering Fable 5-level intelligence at a lower price point.

Key Points
AI Index v4.1.1 includes 26 of 595 models, with Ling 3.0 Tiny, Muse Spark 1.2, Qwen3.8 Max, and Agnes 7 Pro Alpha among the latest additions (Aug 5-29).
Claude Opus 5 leads agentic knowledge work with Fable 5-level intelligence at a lower cost per task.
DeepSeek V4 Flash scores 50 on AI Index, up from its previous score of 40, showcasing improved performance and efficiency.
Inkling Small matches Inkling's score but uses less than a third of the parameters, highlighting advancements in model optimization.
The Cost per Task metric now shows slight absolute increases due to updated methodology, though relative positioning remains unchanged.
Why It Matters
If you're evaluating language models for cost-effective performance, Claude Opus 5 offers Fable 5-level intelligence at a lower price point. For those tracking model efficiency and optimization, Inkling Small's near-equal score to Inkling with fewer parameters is noteworthy.
Frequently Asked Questions
Why does this matter?
If you're evaluating language models for cost-effective performance, Claude Opus 5 offers Fable 5-level intelligence at a lower price point. For those tracking model efficiency and optimization, Inkling Small's near-equal score to Inkling with fewer parameters is noteworthy.
What happened?
The Artificial Analysis Intelligence Index v4.1.1 introduces updates to its methodology and scores new language models like Claude Opus 5 and DeepSeek V4 Flash. The cost per task metric now shows slight absolute increases but maintains relative positioning.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.