01 · When GRPO stalls
When every rollout fails, the model learns nothing useful
Think of GRPO as a harsh grader that only looks at the final score. In Sokoban, 3D navigation, or robot manipulation, one irreversible pivot turn can ruin a long episode, and the algorithm still cannot point to that turn.
Failure mode A · σ = 0
Zero-gradient silence
Sample eight trajectories under a weak policy. If every return matches, group variance is zero. The advantage vanishes, and the model gets no update at all, even when the failures are different mistakes.
- 48.3%
- Sokoban groups frozen
- 16.5%
- PrimitiveSkill groups frozen
Failure mode B · σ > 0
Blunt, collective credit
Even with a −0.1 step penalty, GRPO still paints one scalar onto every action token. It scolds the agent for taking too long, and never isolates the fatal operation at t*.
- σ > 0
- Still episode-level
- t*
- Still unnamed
What the returns look like
Next question OPD/OPSD appears to solve credit assignment failure for GRPO, but do the textual skill hints as the hindsight actually matter?
02 · What the probe reveals
Skill hints help a little. State rollback does most of the work.
Many people assumed OPD/OPSD work because a stronger teacher writes better skill text. We treated that as a hypothesis to test, not a conclusion.
We define the pivot step t* as the first action that leaves the remaining budget infeasible. Then we rewind failed rollouts and resample one suffix from the same base policy, changing only two knobs: when we restore, and whether we append a skill hint.
Pivot step
t* = min { t : Feas(st+1, T − t − 1) = 0 }
- t*
- first action that makes the remaining budget infeasible, else T−1
- Feas(s, k)
- 1 if state s can still reach the goal in k steps
- st+1
- state after action at
- T
- realized trajectory length
Four arms share the same failed starts: no-hint at t*, no-hint at t*+1, 3B-self (self-written skill), and 8B-teacher (Qwen3-VL-8B skill).
Finding 1
State rollback leads
Restore the physical state at t* with no text at all, and suffix return lift ΔR already covers most of the gain (3.3% to 56.6% across tasks). The same restore also cuts zero-variance groups (Sokoban 48.3% to 41.7%, PrimitiveSkill 16.5% to 3.8%), bringing back a usable GRPO signal where the baseline was silent.
Finding 2
Skill text is secondary
Self-written hints differ from no-hint by at most 2.6 points. An 8B teacher adds only 0.3 to 3.5 on top of the 3B self hint. On Navigation, plain no-hint at t* is already best.
Finding 3
One step late collapses recovery
Rewind at t*+1 instead of t*, and ΔR drops hard (Sokoban 3.3% to 0.9%, PrimitiveSkill 17.7% to 8.3%, FrozenLake 18.2% to 8.1%). Timing precision is the key, not prompt polish.
What each probe arm is allowed to see
All arms start from the same failed untrained Qwen2.5-VL-3B rollouts and resample one suffix with that base policy. no-hint restores st* and adds nothing to the prompt. no-hint (t*+1) restores one step later. 3B-self appends failure mode m* and a skill from πbase. 8B-teacher appends a skill from Qwen3-VL-8B.
The catch Recovery wants a state rollback to t*. Real environments and online RL rarely allow one. That is the rollback paradox.
03 · Putting state rollback in the weights
How PIVOT turns state rollback into parameters
The probe leaves a paradox: recovery needs a restore at t*, yet online RL cannot restart the world, and deployment cannot carry an external teacher plus long skill prompts.
Rollback paradox Recovery comes mainly from rolling the state back to the pivot. Explicit physical state rollbacks are too costly online, and impossible in many real settings.
PIVOT’s answer is one VLM πθ that plays three roles in training, then ships as a plain Student at test time: find the pivot without resetting, move credit onto tokens, and drop the extras at deploy.
Stages
Cold-start the Analyzer
Turn each failed rollout into a visual collage plus action log. SFT teaches the model to output the pivot t*, failure mode, and an optional skill. No latent state. About 960 to 1.7k accepted trajectories per environment, then θ←θsft.
Internalize state rollback
Each update samples eight rollouts. On a failure the Analyzer predicts the pivot and crops a three-frame visual panel. A stop-gradient Teacher re-scores the same failed tokens as if that panel were in context. No simulator reset required.
Deploy the Student only
Strip Analyzer and Teacher. The agent runs as πθ on ordinary history: zero extra compute, zero extra parameters, zero skill prompts.
Roles
Reads the full visual collage and action log. Finds the pivot and failure mode without peeking at hidden simulator state.
Re-scores the Student’s failed tokens under a three-frame panel around the predicted pivot. Stop-gradient. Same tokens, as if the state rollback were visible.
Acts on ordinary history. Trains with GRPO plus a confidence-gated OPD term on failed rollouts, so silent GRPO groups can still move.
Joint objective
ℒ = ℒGRPO + λ ℒgated-OPD
Failure modes the Analyzer is asked to name
- Sokoban and FrozenLake: timeout if the goal is still reachable with unlimited steps, deadlock if it is not.
- Navigation: format, blocked, stagnation, or budget.
- PrimitiveSkill: timeout, format, wrong target, or order error.
- SVG: format, regression, stagnation, or mismatch on the last scored canvas.
Inside the recipe · Visual credit
The panel is the internalized state rollback
The Teacher never resets the simulator. A three-frame neighborhood around the predicted pivot is enough to re-score the original failed tokens.
Analyzer SFT examples
Stage I trains the Analyzer on failed rollouts: a visual collage, a compact action log, and a JSON target for pivot, mode, and skill. Only observable history is used. Splits are 90/10; the same recipe applies to the 2B backbone.
| Environment | Accepted | Train | Val |
|---|---|---|---|
| Sokoban | 1170 | 1053 | 117 |
| FrozenLake | 1375 | 1237 | 138 |
| Navigation | 959 | 863 | 96 |
| PrimitiveSkill | 1677 | 1509 | 168 |
| SVG | 1719 | 1547 | 172 |
04 · Results across five tasks
From Sokoban to 3D control, PIVOT sets a new bar
We evaluate on cognitive grids (Sokoban, FrozenLake), embodied 3D control (Navigation, PrimitiveSkill), and generative SVG. Full PIVOT is the M+P+R stack, and it ships with no extra test-time machinery.
| Method | Cognitive Grid | Embodied 3D Control | Generative | All | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Sok. | Frozen Lake |
Navigation | PrimitiveSkill | SVG | ||||||||||
| Base | Com. | Avg. | Place | Stack | Draw. | Align | Avg. | DINO | DS | Avg. | ||||
| Open-source VLMs | ||||||||||||||
| Qwen2.5-VL-72B | 0.20 | 0.44 | 0.70 | 0.77 | 0.74 | 1.00 | 0.50 | 0.00 | 1.00 | 0.63 | 0.84 | 0.62 | 0.73 | 0.55 |
| Qwen2.5-VL-7B | 0.14 | 0.14 | 0.33 | 0.38 | 0.35 | 0.00 | 0.00 | 0.00 | 0.75 | 0.19 | 0.84 | 0.27 | 0.56 | 0.28 |
| Qwen2.5-VL-3B | 0.13 | 0.14 | 0.20 | 0.26 | 0.23 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.79 | 0.30 | 0.54 | 0.21 |
| VLM-R1-3B | 0.16 | 0.15 | 0.33 | 0.34 | 0.34 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.79 | 0.27 | 0.54 | 0.24 |
| Proprietary VLMs | ||||||||||||||
| o4-mini | 0.44 | 0.82 | 0.75 | 0.75 | 0.75 | 1.00 | 0.50 | 0.00 | 0.75 | 0.56 | 0.90 | 0.66 | 0.78 | 0.67 |
| GPT-4o | 0.43 | 0.54 | 0.75 | 0.69 | 0.72 | 0.50 | 0.63 | 0.00 | 0.88 | 0.50 | 0.91 | 0.69 | 0.80 | 0.60 |
| Gemini 2.5 Pro | 0.58 | 0.78 | 0.63 | 0.63 | 0.63 | 0.63 | 0.63 | 0.00 | 0.75 | 0.50 | 0.93 | 0.78 | 0.86 | 0.67 |
| Claude 4.5 Sonnet | 0.31 | 0.80 | 0.67 | 0.67 | 0.67 | 0.63 | 0.50 | 0.00 | 1.00 | 0.53 | 0.95 | 0.81 | 0.88 | 0.64 |
| Claude 3.7 Sonnet | 0.25 | 0.69 | 0.48 | 0.47 | 0.47 | 0.63 | 0.13 | 0.00 | 1.00 | 0.44 | 0.94 | 0.77 | 0.85 | 0.54 |
| Previous SOTA / OPD baseline | ||||||||||||||
| VAGEN-Full | 0.79 | 0.72 | 0.80 | 0.81 | 0.81 | 1.00 | 0.88 | 1.00 | 1.00 | 0.97 | 0.90 | 0.66 | 0.78 | 0.81 |
| GLANCE-Full | 0.85 | 0.78 | 0.86 | 0.88 | 0.87 | 1.00 | 0.88 | 1.00 | 1.00 | 0.97 | 0.92 | 0.70 | 0.81 | 0.86 |
| SDAR | 0.52 | 0.77 | 0.88 | 0.84 | 0.86 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.91 | 0.68 | 0.80 | 0.79 |
| PIVOT (Qwen2.5-VL-3B-Instruct) | ||||||||||||||
| Vanilla-GRPO | 0.78 | 0.71 | 0.87 | 0.75 | 0.81 | 0.92 | 0.86 | 0.89 | 0.93 | 0.90 | 0.90 | 0.66 | 0.78 | 0.80 |
| SFT-GRPO | 0.82 | 0.75 | 0.93 | 0.77 | 0.85 | 0.96 | 0.90 | 0.93 | 0.97 | 0.94 | 0.92 | 0.68 | 0.80 | 0.83 |
| M | 0.86 | 0.83 | 0.94 | 0.78 | 0.86 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.93 | 0.73 | 0.83 | 0.88 |
| M+P | 0.90 | 0.77 | 0.95 | 0.83 | 0.89 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.93 | 0.74 | 0.84 | 0.88 |
| M+P+R | 0.95+16% | 0.84+12% | 0.97+4% | 0.83+8% | 0.90+6% | 1.00+4% | 1.00+11% | 1.00+8% | 1.00+3% | 1.00+6% | 0.93+1% | 0.71+4% | 0.82+3% | 0.90+8% |
Cyan deltas are gains of M+P+R over SFT-GRPO. Enriching Teacher context from M to M+P to M+P+R lifts overall accuracy to 0.90 (+8% over SFT-GRPO, +5% over GLANCE-Full).
| Method | Cognitive Grid | Embodied 3D Control | Generative | All | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Sok. | Frozen Lake |
Navigation | PrimitiveSkill | SVG | ||||||||||
| Base | Com. | Avg. | Place | Stack | Draw. | Align | Avg. | DINO | DS | Avg. | ||||
| Open-source VLMs | ||||||||||||||
| Qwen2.5-VL-72B | 0.20 | 0.44 | 0.70 | 0.77 | 0.74 | 1.00 | 0.50 | 0.00 | 1.00 | 0.63 | 0.84 | 0.62 | 0.73 | 0.55 |
| Qwen2.5-VL-7B | 0.14 | 0.14 | 0.33 | 0.38 | 0.35 | 0.00 | 0.00 | 0.00 | 0.75 | 0.19 | 0.84 | 0.27 | 0.56 | 0.28 |
| Qwen2.5-VL-3B | 0.13 | 0.14 | 0.20 | 0.26 | 0.23 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.79 | 0.30 | 0.54 | 0.21 |
| VLM-R1-3B | 0.16 | 0.15 | 0.33 | 0.34 | 0.34 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.79 | 0.27 | 0.54 | 0.24 |
| Proprietary VLMs | ||||||||||||||
| o4-mini | 0.44 | 0.82 | 0.75 | 0.75 | 0.75 | 1.00 | 0.50 | 0.00 | 0.75 | 0.56 | 0.90 | 0.66 | 0.78 | 0.67 |
| GPT-4o | 0.43 | 0.54 | 0.75 | 0.69 | 0.72 | 0.50 | 0.63 | 0.00 | 0.88 | 0.50 | 0.91 | 0.69 | 0.80 | 0.60 |
| Gemini 2.5 Pro | 0.58 | 0.78 | 0.63 | 0.63 | 0.63 | 0.63 | 0.63 | 0.00 | 0.75 | 0.50 | 0.93 | 0.78 | 0.86 | 0.67 |
| Claude 4.5 Sonnet | 0.31 | 0.80 | 0.67 | 0.67 | 0.67 | 0.63 | 0.50 | 0.00 | 1.00 | 0.53 | 0.95 | 0.81 | 0.88 | 0.64 |
| Claude 3.7 Sonnet | 0.25 | 0.69 | 0.48 | 0.47 | 0.47 | 0.63 | 0.13 | 0.00 | 1.00 | 0.44 | 0.94 | 0.77 | 0.85 | 0.54 |
| Previous SOTA / OPD baseline | ||||||||||||||
| VAGEN-Full | 0.79 | 0.72 | 0.80 | 0.81 | 0.81 | 1.00 | 0.88 | 1.00 | 1.00 | 0.97 | 0.90 | 0.66 | 0.78 | 0.81 |
| GLANCE-Full | 0.85 | 0.78 | 0.86 | 0.88 | 0.87 | 1.00 | 0.88 | 1.00 | 1.00 | 0.97 | 0.92 | 0.70 | 0.81 | 0.86 |
| SDAR | 0.52 | 0.77 | 0.88 | 0.84 | 0.86 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.91 | 0.68 | 0.80 | 0.79 |
| PIVOT (Qwen3-VL-2B-Instruct) | ||||||||||||||
| Vanilla-GRPO | 0.77 | 0.89 | 0.97 | 0.73 | 0.85 | 0.77 | 0.77 | 0.77 | 0.77 | 0.77 | 0.90 | 0.68 | 0.79 | 0.81 |
| SFT-GRPO | 0.80 | 0.87 | 0.92 | 0.74 | 0.83 | 0.83 | 0.80 | 0.81 | 0.84 | 0.82 | 0.91 | 0.67 | 0.79 | 0.82 |
| M | 0.82 | 0.81 | 0.93 | 0.75 | 0.84 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.92 | 0.70 | 0.81 | 0.86 |
| M+P | 0.97 | 0.89 | 0.97 | 0.80 | 0.89 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.93 | 0.71 | 0.82 | 0.91 |
| M+P+R | 0.92+15% | 0.90+3% | 0.98+7% | 0.91+23% | 0.95+14% | 1.00+20% | 1.00+25% | 1.00+23% | 1.00+19% | 1.00+22% | 0.93+2% | 0.76+13% | 0.85+8% | 0.92+12% |
Same baselines; PIVOT rows use Qwen3-VL-2B. Privileged context scales from M (0.86) to M+P (0.91) to M+P+R (0.92, +12% over SFT-GRPO). Sokoban peaks at M+P (0.97).
Ablation: localization is not optional
A panel alone or a failure mode alone already helps. Combining them raises Sokoban / Navigation to about 0.90 / 0.89, and full M+P+R peaks at 0.95 / 0.90.
Replace the localized step with a random index (M+P-Random+R) and scores fall to 0.75 / 0.80, below Vanilla-GRPO. Coarse context is noise. Precise pivot localization is the ingredient that matters.
| Variant | Sok. | Nav. |
|---|---|---|
| Vanilla-GRPO | 0.78 | 0.81 |
| P | 0.86 | 0.85 |
| M | 0.86 | 0.86 |
| M+P | 0.90 | 0.89 |
| M+P+R | 0.95 | 0.90 |
| M+P-Random+R | 0.75 | 0.80 |
Training gets cleaner, not slower
Self-evolving pivot localization
On FrozenLake we can check Analyzer accuracy against a shortest-path oracle. Under M+P+R it climbs from about 0.56 toward 0.95 as training proceeds. Better localization feeds better credit, without a separate diagnostic model.
Training efficiency
PIVOT kills bad actions earlier, so failed rollouts shorten. On Sokoban each RL step takes 95s versus 102s for Vanilla-GRPO, and far below a frozen 8B teacher at 373s. At test time the diagnostic branches are gone, so there is no parameter or latency tax.