vLLM V0 to V1: Correctness Before Corrections in RL

ServiceNow details how its PipelineRL team moved RL rollout generation from vLLM V0 to V1 by fixing four backend gaps — logprob semantics, runtime defaults, inflight weight updates, and an fp32 lm_head — before changing the objective.

Wednesday May 6, 2026 Source: huggingface.co
TL;DR — Quick Answer

ServiceNow's PipelineRL team migrated RL rollout generation from vLLM V0 (0.8.5) to vLLM V1 (0.18.1) and restored trainer-side parity by fixing four backend issues before touching the objective. The fixes: setting logprobs-mode to processed_logprobs so V1 returns logprobs from the processed sampling distribution, disabling V1-only runtime defaults like prefix caching and async scheduling, matching the V0 inflight weight-update behavior with pause_generation mode keep and clear_cache false, and computing the final projection with an fp32 lm_head. After the fixes, the corrected V1 run tracked the V0 reference across clip rate, KL, entropy, and reward. The team's conclusion: fix inference backend correctness first, then add objective-side corrections for whatever mismatch remains.

Key Takeaways

vLLM V0 to V1: Correctness Before Corrections in RL — AI news article illustration

ServiceNow’s PipelineRL teams train reinforcement learning agents with vLLM as the inference engine for rollout generation. Those rollouts produce the logprobs that compute policy ratios, KL, clip rate, entropy, and reward — so when a migration from vLLM V0 to V1 split the training curves, the source of the mismatch became an engineering problem worth an entire postmortem.

The Migration Target

The migration goal was deliberately narrow: verify V1 returns rollout logprobs in the form the trainer expects, rerun the same workload against the V0 reference, and only then evaluate objective-level changes.

The first symptoms showed up in clamp_log_ratio_new_old_indicator, KL, entropy, and reward from a GSPO training run — the same class of mismatch that can surface in PPO, GRPO, or any online RL system that treats rollout-side logprobs as part of the optimization target.

Failure Modes

The team separated possible causes into three layers:

  1. Semantic mismatch — the backend returns logprobs with different meaning than the trainer expects
  2. Inference-path mismatch — different runtime defaults send the same prompts down a different execution path
  3. Objective mismatch — the RL objective needs correction for stale or off-policy rollouts

The useful diagnosis came from ruling out the first two as backend behavior problems before blaming the objective.

Four Backend Fixes

The Remaining Gap: fp32 lm_head

Even after the backend fixes, full parity required matching the numerical path for logits. The trainer uses an fp32 lm_head for the final projection, and the rollout backend had to match. The same issue appears in the MiniMax-M1 technical report, which traced a training-inference token-probability mismatch to the LM output head and fixed it by computing the head in fp32. The ScaleRL paper later includes fp32 logits and head computation as part of its RL recipe.

The Lesson: Backend First

Ablations showed why each fix was necessary: processed_logprobs alone fixed the semantic bug but left the training mismatch; the first V1 attempt was a confounded comparison because multiple V1-only defaults were enabled; batch invariance did not reproduce parity under higher lag and NCCL complications.

The takeaway is narrow but firm: fix backend correctness first, then add corrections for the mismatch that remains. Only after inference parity — via the four fixes — did the final V1 run track the V0 reference across clip rate, KL, entropy, and reward.

Frequently Asked Questions

What is PipelineRL?

PipelineRL is ServiceNow's reinforcement learning pipeline for training agents and models, which uses the vLLM inference engine to generate rollouts and consume the resulting token logprobs for policy updates.

Why did ServiceNow migrate from vLLM V0 to V1?

vLLM V1 is a rewritten, faster engine, but rollout-side logprobs and runtime behavior differed from V0, causing a train-inference mismatch that broke RL metrics like clip rate, KL, entropy, and reward.

What fixed the vLLM V1 training mismatch?

Four backend fixes: processed_logprobs mode, disabling V1 default prefix caching and async scheduling, matching V0's inflight weight-update behavior, and computing the lm_head in fp32.

Why fix the backend before the RL objective?

Objective-side corrections like truncated importance sampling are useful, but mixing them with a broken inference backend makes training curves hard to interpret — backend correctness must come first.

This article is based on the official announcement from huggingface.co . Read the original for full technical details.

Related Articles

Back to all news