Nous Research released NousCoder-14B on January 7, 2026, a competitive programming model that the company says matches or beats several larger proprietary systems. It arrived during a surge of excitement around Anthropic’s Claude Code, sharpening the contrast between closed agentic tools and reproducible open models.
What NousCoder-14B Achieves
The 14B model scores 67.87 percent on LiveCodeBench v6, a standardized evaluation of competitive programming problems, up 7.08 percentage points over its Alibaba Qwen3-14B base. It is released under an Apache 2.0 license.
Trained in Four Days
Training took only about four days on 48 Nvidia B200 GPUs. The pipeline used verifiable rewards, executing generated code against test cases and returning a binary pass or fail, with Modal handling sandboxed parallel execution under 15 second and 4 gigabyte limits.
- 24,000 competitive programming problems, each with hundreds of test cases
- DAPO for reinforcement learning, chosen after comparison with alternatives
- Iterative context extension, from 32,000 to 40,000 tokens, with evaluation at about 80,000
- Overlapped inference and verification to keep expensive GPUs busy
The Open Recipe
What sets the release apart is its openness. Nous Research published the weights, the complete Atropos reinforcement learning environment, the benchmark suite and the training harness, so other researchers can reproduce or extend the work.
The Data Wall
A striking finding in the technical report is that the 24,000 problems represent a significant share of all readily available, verifiable competitive programming problems. The author, Joe Li, concludes the domain is approaching the limits of high-quality data, pointing to synthetic data generation and self-play as the next frontier.
What Comes Next
Li estimates the model’s jump mirrors his own climb on Codeforces from a 1600 to 1750 rating to 2100 to 2200, a leap that took him two years and 1,000 problems. The model needed 24,000. The next challenge is teaching models to write their own problems.