TL;DR
AWS expanded the context window for OpenAI’s GPT-5.6 Sol, Terra, and Luna models on Bedrock to 1 million tokens, enabling single-inference processing of entire code repositories. Prompt caching offers a 90% discount on repeated context, and AWS announced up to 80% price reductions aligning with OpenAI’s first-party rates.
What’s New
AWS Bedrock updates for GPT-5.6 models:
- 1M-token context: Sol, Terra, and Luna all support 1M tokens
- Single-inference codebases: Entire code repositories processed in one call
- Prompt caching: 90% discount on repeated context
- Price cuts: Up to 80% reduction, matching OpenAI’s first-party pricing
The updates make GPT-5.6 models more practical for enterprise workloads on AWS infrastructure.
What It Means
- Whole-repository analysis: Developers can analyze entire codebases in one inference
- Long documents: Legal, research, and documentation workloads at scale
- Cost efficiency: Prompt caching dramatically reduces repeated-context costs
- Enterprise access: Bedrock gives enterprises AWS-native access to GPT-5.6
Context
The AWS update coincides with the broader price war:
- OpenAI direct: Luna cut 80% to $0.20/M input
- DeepSeek: $0.14/M input — even cheaper
- Alibaba: Qwen3.8-Max at $2/M input
Bedrock’s role: AWS hosts models from OpenAI, Anthropic, Meta, Mistral, and others — betting platform breadth beats owning the best model.
Implications
- Enterprise adoption: Lower prices and longer context accelerate enterprise use
- Multi-model strategy: AWS benefits from hosting all major models
- Competition: Azure and Google Cloud must match Bedrock’s offerings
- Amazon’s retreat: The Nova model pullback makes AWS more dependent on partners
For the AI industry, GPT-5.6 on Bedrock with 1M-token context and 80% price cuts demonstrates how quickly frontier capability is becoming cheap and accessible through cloud platforms.