Claude Fable 5.1 (Max) by @AnthropicAI has landed in the Agent Arena at #1 with +15.8% net improveme...
lmarena.ai(@lmarena_ai) · 人工智能
Claude Fable 5.1 (Max) by @AnthropicAI has landed in the Agent Arena at #1 with +15.8% net improvement across 6.7k+ real-world agentic sessions! It also redraws the price-performance frontier: #1 on the leaderboard at a median cost of $4.14/task. By signal, Claude Fable 5.1 sees a massive lead in implicit user sentiment with an astonishing (+42.5%) in Praise vs. Complaint. Users are praising it around 2x more often than the next top model. It also sees strong explicit feedback via Confirmed Success (+22.4%), and solid Bash Recovery (+13.1%), with no Tool Hallucinations. More detail on its performance by signal below. Claude Fable 5.1 (Max) out ranks all past Claude variants and the rest of the pack by a healthy lead. In Agent Arena, we measure models on millions of long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. Stay tuned as traces continue to come in for the latest GPT-6 Astra to see how it compares. Use Agent Mode to contribute to the real-world rankings. Congrats again to @AnthropicAI for this release. Claude @claudeai We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work. Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 15 🔄 12 ❤️ 183 👀 31174 📊 29 ⚡ Powered by xgo.ing