ServiceNow’s PipelineRL teams train reinforcement learning agents with vLLM as the inference engine for rollout generation. Those rollouts produce the logprobs that compute policy ratios, KL, clip rate, entropy, and reward — so when a migration from vLLM V0 to V1 split the training curves, the source of the mismatch became an engineering problem worth an entire postmortem.
The Migration Target
The migration goal was deliberately narrow: verify V1 returns rollout logprobs in the form the trainer expects, rerun the same workload against the V0 reference, and only then evaluate objective-level changes.
The first symptoms showed up in clamp_log_ratio_new_old_indicator, KL, entropy, and reward from a GSPO training run — the same class of mismatch that can surface in PPO, GRPO, or any online RL system that treats rollout-side logprobs as part of the optimization target.
Failure Modes
The team separated possible causes into three layers:
- Semantic mismatch — the backend returns logprobs with different meaning than the trainer expects
- Inference-path mismatch — different runtime defaults send the same prompts down a different execution path
- Objective mismatch — the RL objective needs correction for stale or off-policy rollouts
The useful diagnosis came from ruling out the first two as backend behavior problems before blaming the objective.
Four Backend Fixes
- Logprob semantics — V1 returns raw logprobs before logits post-processing by default; setting
logprobs-mode=processed_logprobsremoved the mean offset - Runtime defaults — V1-only defaults for prefix caching and async scheduling were disabled to match the V0 path, since a prefix-cache hit can reuse state computed before a weight update
- Inflight weight updates — the V1 analogue of V0’s update model was
pause_generation(mode=keep, clear_cache=False)plus a collective RPC weight update, preserving cached state like V0 - fp32 lm_head — the final projection must match the trainer’s fp32 head
The Remaining Gap: fp32 lm_head
Even after the backend fixes, full parity required matching the numerical path for logits. The trainer uses an fp32 lm_head for the final projection, and the rollout backend had to match. The same issue appears in the MiniMax-M1 technical report, which traced a training-inference token-probability mismatch to the LM output head and fixed it by computing the head in fp32. The ScaleRL paper later includes fp32 logits and head computation as part of its RL recipe.
The Lesson: Backend First
Ablations showed why each fix was necessary: processed_logprobs alone fixed the semantic bug but left the training mismatch; the first V1 attempt was a confounded comparison because multiple V1-only defaults were enabled; batch invariance did not reproduce parity under higher lag and NCCL complications.
The takeaway is narrow but firm: fix backend correctness first, then add corrections for the mismatch that remains. Only after inference parity — via the four fixes — did the final V1 run track the V0 reference across clip rate, KL, entropy, and reward.