3 articles about evaluation
New research examines whether LLM benchmark scores reflect real model capability, revealing how measurement artifacts and benchmark design choices distort our picture of model performance.
An open benchmark from IBM Research that compares full AI agent systems — not just models — across six tasks, reporting both quality and cost per task.
To fight test-set contamination, the Open ASR Leaderboard adds private evaluation datasets from Appen and DataoceanAI, keeping the default average WER on public sets while exposing optional private metrics.