FILTER: BENCHMARKS
Blog
OpenAIAILLMBenchmarks
GPT-5.6's Three Tiers: Sol, Terra, and Luna, Simply Explained
OpenAI split GPT-5.5 into three models instead of shipping one bigger one. Simply explained: what Sol, Terra, and Luna actually are, what they cost, what's genuinely new, and which one you should reach for.
July 2, 2026
6 min read
LLMAIEvalsBenchmarksEngineering
The State of AI Benchmarks in 2026
Classic benchmarks are saturated, contaminated, and increasingly useless for choosing a model. A practitioner's guide to what frontier evals actually measure, why leaderboards lie, and how to build the evals that matter for your specific use case.
April 3, 2026
18 min read