For the past three years, enterprise AI procurement has run on a stable assumption: the safest choice is the largest frontier model available, with smaller models reserved for workloads that tolerate quality loss in exchange for lower cost. A new community article from Dharma-AI on Hugging Face argues that comparison set is incomplete.
The Strategic Default
The procurement default exists because it was mostly correct. When GPT-4 launched, it beat every smaller model on the benchmarks that mattered, and the pattern repeated through Claude 3, Gemini 1.5, and each 2025 frontier generation. Following OpenAI’s scaling laws, capability appeared to scale with parameters and training compute, so defaulting to scale was the rational move. What the default missed, Dharma argues, is a different kind of model: not a smaller frontier model, but a specialized one whose training history has been deliberately moved toward the task.
The DharmaOCR Result
In April 2026, Dharma released DharmaOCR, a pair of specialized small language models for structured OCR, alongside a benchmark and a companion paper (arXiv 2604.14314). On extraction quality, the specialized 3-billion-parameter model scored 0.911, ahead of every frontier API tested:
- Claude Opus 4.6 at 0.833
- Gemini 3.1 Pro at 0.820
- GPT-5.4 at 0.750
- Google Vision at 0.686, Google Document AI at 0.640, Amazon Textract at 0.618
On cost, the gap was far wider: the specialized 3B model ran at roughly fifty-two times lower cost per million pages than Claude Opus 4.6. And on production stability, it posted the lowest text-degeneration rate evaluated, at 0.20%.
Why Specialization Won
The paper names the variable directly: how close a model’s training trajectory has been moved to its deployment task, not parameter count. The evidence: Nanonets-OCR2-3B, already specialized for general OCR, fine-tuned on the target domain reached 0.921 with a 0.20% degeneration rate. Qwen2.5-VL-3B, an identical-architecture general-purpose model run through the same pipeline, reached just 0.793 with 1.41% degeneration. Same architecture, same training, different starting position.
Specialization Compounds
Dharma’s data shows alignment behaves like a hierarchy rather than a binary state. At the 7B scale, the same training lifted olmOCR-2-7B to 0.927 with 0.40% degeneration versus 0.906 and 1.01% for the general-purpose Qwen2.5-VL-7B start. Starting closer to the task compounds the benefit of each downstream fine-tuning stage. Buyers, the article argues, should elevate distributional alignment alongside parameter count as a first-class evaluation variable - and question whether frontier benchmark leadership alone is sufficient evidence for procurement decisions.