#evaluation

3 articles about evaluation

BenchMIRT Asks: What Are LLM Benchmarks Actually Measuring?
AI Research Sep 1, 2026

BenchMIRT Asks: What Are LLM Benchmarks Actually Measuring?

New research examines whether LLM benchmark scores reflect real model capability, revealing how measurement artifacts and benchmark design choices distort our picture of model performance.

The Open Agent Leaderboard
AI Research May 18, 2026

The Open Agent Leaderboard

An open benchmark from IBM Research that compares full AI agent systems — not just models — across six tasks, reporting both quality and cost per task.

Adding Benchmaxxer Repellant to the Open ASR Leaderboard
AI Research May 6, 2026

Adding Benchmaxxer Repellant to the Open ASR Leaderboard

To fight test-set contamination, the Open ASR Leaderboard adds private evaluation datasets from Appen and DataoceanAI, keeping the default average WER on public sets while exposing optional private metrics.