#benchmarks

8 articles about benchmarks

GPT-6 Astra: OpenAI Claims New Intelligence Generation — 98% FrontierMath, Human Parity on ARC-AGI-3, and New Prime Number Proofs
AI Models Sep 3, 2026

GPT-6 Astra: OpenAI Claims New Intelligence Generation — 98% FrontierMath, Human Parity on ARC-AGI-3, and New Prime Number Proofs

OpenAI's GPT-6 Astra capabilities announcement: state-of-the-art on computer use, coding, and science with saturated benchmarks, two real prime-gap proofs, and 0% rate of escaping authorized scope vs Sol's 48%.

BenchMIRT Asks: What Are LLM Benchmarks Actually Measuring?
AI Research Sep 1, 2026

BenchMIRT Asks: What Are LLM Benchmarks Actually Measuring?

New research examines whether LLM benchmark scores reflect real model capability, revealing how measurement artifacts and benchmark design choices distort our picture of model performance.

Anthropic Releases Claude Opus 5: Matches Fable 5 Performance at Half the Price
AI Research Jul 24, 2026

Anthropic Releases Claude Opus 5: Matches Fable 5 Performance at Half the Price

Anthropic launched Claude Opus 5, scoring 43.3% on Frontier-Bench v0.1 — surpassing all competitors. Priced at $5/M input and $25/M output tokens, it delivers near-Fable 5 intelligence at half the cost per task, making it the new default on Claude Max.

Anthropic Launches Claude Sonnet 5: First Sonnet-Class Model to Outscore Opus 4.8
AI Models Jun 30, 2026

Anthropic Launches Claude Sonnet 5: First Sonnet-Class Model to Outscore Opus 4.8

Anthropic released Claude Sonnet 5 on June 30, 2026, making it the default model for Free and Pro plans. Priced at $2 per million input tokens, it is the first Sonnet-class model to outscore the Opus 4.8 flagship.

Grok Imagine Video 1.5 Beats Sora, Veo, and Kling in Blind User Benchmarks at 86% Lower Cost
AI Tools Jun 22, 2026

Grok Imagine Video 1.5 Beats Sora, Veo, and Kling in Blind User Benchmarks at 86% Lower Cost

Grok Imagine Video 1.5 beat Sora 2, Veo 3.1, and Kling in blind user benchmarks at 86% lower cost, with native audio and dialogue in one pass.

The Open Agent Leaderboard
AI Research May 18, 2026

The Open Agent Leaderboard

An open benchmark from IBM Research that compares full AI agent systems — not just models — across six tasks, reporting both quality and cost per task.

DeepSeek V4 Outperforms Claude Opus 4.6 in Benchmarks
AI Research Apr 26, 2026

DeepSeek V4 Outperforms Claude Opus 4.6 in Benchmarks

Chinese AI company DeepSeek launches V4 model with superior efficiency, available under MIT License.

Google DeepMind's Gemini 2.0 Beats GPT-4 on Key Benchmarks
AI Research Apr 17, 2025

Google DeepMind's Gemini 2.0 Beats GPT-4 on Key Benchmarks

Google's Gemini 2.0 Ultra outperforms GPT-4 on reasoning, coding, and multimodal tasks in independent evaluations.