Frontier LLM Meta-Benchmarks
August 2026
Multi-Domain Composite Evaluation across arena.ai, SWE-Pro, Autonomous OSWorld, and Unit Economics
Export CSV
Download HTML
← Back to Main
All Frontier Models
US Frontier (Anthropic / OpenAI / xAI / Google)
China Ecosystem (Qwen / Kimi / DeepSeek / GLM)
European Sovereign (Mistral AI)
Open Weights / Self-Hosted
Arena.ai Code Elo Ratings
Arena.ai Live
SWE-Pro / DeepSWE Accuracy (%)
Software Engineering
Meta-Benchmark Matrix (arena.ai + SWE + Agentic)
Model & Family
Arena.ai Agent Rank
Arena Code Elo
SWE-Pro / DeepSWE
OSWorld / Agents
Context
Input / Output ($/1M)
Primary Architectural Profile