Direct Preference Optimization has been famous for aligning chatbots to human judgments. A new study from Dharma-AI shows the method has a second life: applied to structured OCR, a DPO stage turned the model’s own failures into training signal and cut a production failure mode by more than half — no chatbot required.
The Failure Mode SFT Misses
Text degeneration — a model entering a repetition loop instead of producing a clean transcription — is a signature failure of structured generation. In the DharmaOCR benchmark, vanilla degeneration rates ranged from below 1 percent to above 33 percent across open-source model families. Supervised fine-tuning reduced the problem for most models but rarely to production-acceptable levels, because SFT trains token by token and never penalizes the loop as a completed-output failure.
Degenerate Outputs as Training Signal
The paper’s central design decision inverts standard practice. Where most pipelines filter degenerate outputs out of the dataset as low-quality noise, DharmaOCR kept them as the rejected examples in each chosen-versus-rejected pair. The pipeline:
- Generated candidate outputs from the SFT model across 23,726 documents
- Scored each with an automated LLM judge against four task-specific criteria
- Labeled degeneration loops as rejected and clean extractions as chosen
- Trained a DPO stage on the resulting pairs — a one-time investment after SFT
The authors describe the result as preference-guided implicit unlikelihood: the model is steered toward good output and simultaneously away from a specific class of failure.
Consistent Across Five Model Families
The DPO stage reduced degeneration in every family tested, with the direction invariant regardless of architecture, scale, or starting rate:
- Average reduction of 59.4 percent; best case 87.6 percent (Nanonets-OCR2-3B from 1.61 percent to 0.20 percent)
- gemma-3-4b-it entered with the worst vanilla rate by an order of magnitude — 33.96 percent — and still saw a ~75 percent reduction
- Qwen2.5-VL-3B rose from 0.60 to 3.23 percent after SFT as it gained capability, then DPO brought it to 1.41 percent — evidence that capability and degeneration resistance move independently
When DPO Transfers
The method generalizes where three conditions hold: failures are categorically identifiable (not just slightly worse output), a scorer can reliably separate them from acceptable output without human annotation, and there is enough volume to build a meaningful preference dataset. Any task whose failures are this legible can reuse the same pipeline.
What It Means
For ML engineers, the takeaway is structural: SFT and DPO address different failure dimensions, and a DPO stage is a cheap post-processing investment for structured generation reliability. For the field, it is evidence that preference optimization is a general training technique — not a chatbot trick.