8 articles about benchmarks
OpenAI's GPT-6 Astra capabilities announcement: state-of-the-art on computer use, coding, and science with saturated benchmarks, two real prime-gap proofs, and 0% rate of escaping authorized scope vs Sol's 48%.
New research examines whether LLM benchmark scores reflect real model capability, revealing how measurement artifacts and benchmark design choices distort our picture of model performance.
Anthropic launched Claude Opus 5, scoring 43.3% on Frontier-Bench v0.1 — surpassing all competitors. Priced at $5/M input and $25/M output tokens, it delivers near-Fable 5 intelligence at half the cost per task, making it the new default on Claude Max.
Anthropic released Claude Sonnet 5 on June 30, 2026, making it the default model for Free and Pro plans. Priced at $2 per million input tokens, it is the first Sonnet-class model to outscore the Opus 4.8 flagship.
Grok Imagine Video 1.5 beat Sora 2, Veo 3.1, and Kling in blind user benchmarks at 86% lower cost, with native audio and dialogue in one pass.
An open benchmark from IBM Research that compares full AI agent systems — not just models — across six tasks, reporting both quality and cost per task.
Chinese AI company DeepSeek launches V4 model with superior efficiency, available under MIT License.
Google's Gemini 2.0 Ultra outperforms GPT-4 on reasoning, coding, and multimodal tasks in independent evaluations.