SpaceXAI has launched Grok 4.5, its smartest model yet, purpose-built for coding, agentic tasks, and knowledge work. The release pairs state-of-the-art engineering benchmark scores with pricing that undercuts the frontrunners by roughly half.
The Model
Grok 4.5 was trained across tens of thousands of NVIDIA GB300 GPUs, with heavy investment in data filtering and curation — deduplication, quality scoring, and domain-focused selection — plus scaled reinforcement learning over hundreds of thousands of tasks centered on multi-step software engineering. The asynchronous training stack lets agentic rollouts run for hours while learning continues across the cluster.
Benchmark Results
The numbers land squarely in frontier territory:
- DeepSWE 1.0 (pass@1): 62.0% — behind GPT 5.5 xhigh at 64.31% but ahead of Opus 4.8 max at 55.75%
- SWE Marathon: 29.0% — the top resolve rate, above Opus 4.8 max at 26.0%
- SWE Bench Pro: 64.7% resolve rate against Fable max at 80.4%
- Terminal Bench 2.1: 83.3%
Speed and Token Efficiency
Grok 4.5 is served at fast-model speeds of around 80 tokens per second. It is also markedly more token-efficient, resolving SWE Bench Pro tasks with about 4.2x fewer output tokens than Opus 4.8 max — 15,954 versus 67,020 on average — cutting both latency and cost per completed task.
Pricing
Grok 4.5 is priced at $2 per million input tokens and $6 per million output tokens, roughly half the price of the leading competitors for comparable class performance. Combined with token efficiency, the company says it delivers the highest intelligence per unit of time and cost.
Availability
- Grok Build — now the default model, including Excel, PowerPoint, and Word plugins
- Cursor — available on all plans
- SpaceXAI console — via API, with free usage for a limited time
The pricing structure gives developers a clear reason to switch, and the benchmark table gives them cover to do so.