Rethinking OPSD/OPD: Visual Agents Need Precise State Rollback, Not Fancy Prompts

arXiv Code (Coming Soon) HF Models (Coming Soon) Project page

Pivot step. Across Sokoban, FrozenLake, Navigation, and PrimitiveSkill, the first unrecoverable action (deadlock, blocked path, or wrong target) ends the feasible region.

The whole story

  1. 01

    GRPO grades the whole paper, not the wrong step. When every sample fails, the model goes silent; when a signal remains, it still cannot name the pivot turn.

  2. 02

    OPD/OPSD look like better prompting. Probes show the lift comes mainly from a state rollback to t*. Skill text is a smaller bonus, and one step late collapses recovery.

  3. 03

    PIVOT puts that state rollback inside the weights: one model diagnoses, re-scores, and acts. At test time only the Student remains.

  4. 04

    On five visual agent tasks, full PIVOT reaches 0.90 / 0.92, beats prior SOTA, and trains faster than Vanilla-GRPO while skipping a heavy external teacher.

01 · When GRPO stalls

When every rollout fails, the model learns nothing useful

Think of GRPO as a harsh grader that only looks at the final score. In Sokoban, 3D navigation, or robot manipulation, one irreversible pivot turn can ruin a long episode, and the algorithm still cannot point to that turn.

Failure mode A · σ = 0

Zero-gradient silence

Sample eight trajectories under a weak policy. If every return matches, group variance is zero. The advantage vanishes, and the model gets no update at all, even when the failures are different mistakes.

48.3%
Sokoban groups frozen
16.5%
PrimitiveSkill groups frozen

Untrained Qwen2.5-VL-3B

Failure mode B · σ > 0

Blunt, collective credit

Even with a −0.1 step penalty, GRPO still paints one scalar onto every action token. It scolds the agent for taking too long, and never isolates the fatal operation at t*.

σ > 0
Still episode-level
t*
Still unnamed

Length penalty is not pivot diagnosis

What the returns look like
Red · σ = 0 · silence Blue · σ > 0 · blunt credit
Failed-rollout return distributions. Red marks zero-gradient silence (σ=0); blue marks coarse episode-level credit (σ>0).
Failed return distributions across benchmarks. Red groups give GRPO no signal. Blue groups still only score the whole episode.
Next question OPD/OPSD appears to solve credit assignment failure for GRPO, but do the textual skill hints as the hindsight actually matter?

02 · What the probe reveals

Skill hints help a little. State rollback does most of the work.

Many people assumed OPD/OPSD work because a stronger teacher writes better skill text. We treated that as a hypothesis to test, not a conclusion.

We define the pivot step t* as the first action that leaves the remaining budget infeasible. Then we rewind failed rollouts and resample one suffix from the same base policy, changing only two knobs: when we restore, and whether we append a skill hint.

Pivot step

t* = min { t : Feas(st+1, T − t − 1) = 0 }

t*
first action that makes the remaining budget infeasible, else T−1
Feas(s, k)
1 if state s can still reach the goal in k steps
st+1
state after action at
T
realized trajectory length

Four arms share the same failed starts: no-hint at t*, no-hint at t*+1, 3B-self (self-written skill), and 8B-teacher (Qwen3-VL-8B skill).

Finding 1

State rollback leads

Restore the physical state at t* with no text at all, and suffix return lift ΔR already covers most of the gain (3.3% to 56.6% across tasks). The same restore also cuts zero-variance groups (Sokoban 48.3% to 41.7%, PrimitiveSkill 16.5% to 3.8%), bringing back a usable GRPO signal where the baseline was silent.

Finding 2

Skill text is secondary

Self-written hints differ from no-hint by at most 2.6 points. An 8B teacher adds only 0.3 to 3.5 on top of the 3B self hint. On Navigation, plain no-hint at t* is already best.

Finding 3

One step late collapses recovery

Rewind at t*+1 instead of t*, and ΔR drops hard (Sokoban 3.3% to 0.9%, PrimitiveSkill 17.7% to 8.3%, FrozenLake 18.2% to 8.1%). Timing precision is the key, not prompt polish.

Suffix-return lift rate under four rollback probes across five tasks.
Suffix return lift ΔR (%). Same failed starts, one resampled suffix, four probe arms.
Zero-variance group rate on Sokoban and PrimitiveSkill under baseline and four rollback probes.
Zero-variance group rate (%). Baseline versus the four rollback probes.
What each probe arm is allowed to see

All arms start from the same failed untrained Qwen2.5-VL-3B rollouts and resample one suffix with that base policy. no-hint restores st* and adds nothing to the prompt. no-hint (t*+1) restores one step later. 3B-self appends failure mode m* and a skill from πbase. 8B-teacher appends a skill from Qwen3-VL-8B.

The catch Recovery wants a state rollback to t*. Real environments and online RL rarely allow one. That is the rollback paradox.

03 · Putting state rollback in the weights

How PIVOT turns state rollback into parameters

The probe leaves a paradox: recovery needs a restore at t*, yet online RL cannot restart the world, and deployment cannot carry an external teacher plus long skill prompts.

Rollback paradox Recovery comes mainly from rolling the state back to the pivot. Explicit physical state rollbacks are too costly online, and impossible in many real settings.

PIVOT’s answer is one VLM πθ that plays three roles in training, then ships as a plain Student at test time: find the pivot without resetting, move credit onto tokens, and drop the extras at deploy.

PIVOT architecture: Stage I analyzer cold-start, Stage II three unified roles, and zero-overhead inference.
Stage I cold-starts the Analyzer on collages and action logs. Stage II trains on the original failed rollouts under a privileged panel. Inference keeps only the Student.

Stages

I

Cold-start the Analyzer

Turn each failed rollout into a visual collage plus action log. SFT teaches the model to output the pivot t*, failure mode, and an optional skill. No latent state. About 960 to 1.7k accepted trajectories per environment, then θ←θsft.

II

Internalize state rollback

Each update samples eight rollouts. On a failure the Analyzer predicts the pivot and crops a three-frame visual panel. A stop-gradient Teacher re-scores the same failed tokens as if that panel were in context. No simulator reset required.

∅

Deploy the Student only

Strip Analyzer and Teacher. The agent runs as πθ on ordinary history: zero extra compute, zero extra parameters, zero skill prompts.

Roles

Analyzer

Reads the full visual collage and action log. Finds the pivot and failure mode without peeking at hidden simulator state.

Privileged Teacher

Re-scores the Student’s failed tokens under a three-frame panel around the predicted pivot. Stop-gradient. Same tokens, as if the state rollback were visible.

Student

Acts on ordinary history. Trains with GRPO plus a confidence-gated OPD term on failed rollouts, so silent GRPO groups can still move.

Joint objective

ℒ = ℒGRPO + λ ℒgated-OPD

λ = 0.01. Confidence-gated OPD (Agarwal et al., 2024; Lu et al., 2026) applies only to failed trajectories, with a sigmoid gate (β = 5). When GRPO is silent, this term can still move the policy. Training runs for 250 updates.
Failure modes the Analyzer is asked to name
  • Sokoban and FrozenLake: timeout if the goal is still reachable with unlimited steps, deadlock if it is not.
  • Navigation: format, blocked, stagnation, or budget.
  • PrimitiveSkill: timeout, format, wrong target, or order error.
  • SVG: format, regression, stagnation, or mismatch on the last scored canvas.

Inside the recipe · Visual credit

The panel is the internalized state rollback

The Teacher never resets the simulator. A three-frame neighborhood around the predicted pivot is enough to re-score the original failed tokens.

Three Sokoban frames around a deadlock pivot.
Sokoban, t* = 4, deadlock.

Analyzer SFT examples

Stage I trains the Analyzer on failed rollouts: a visual collage, a compact action log, and a JSON target for pivot, mode, and skill. Only observable history is used. Splits are 90/10; the same recipe applies to the 2B backbone.

Analyzer SFT splits (3B packs; same recipe on 2B)
Environment Accepted Train Val
Sokoban11701053117
FrozenLake13751237138
Navigation95986396
PrimitiveSkill16771509168
SVG17191547172
Sokoban Analyzer SFT example: trajectory collage, action log, and JSON target.
Sokoban. Collage and action log (left) with JSON supervision (right). Pivot line highlighted in the log.

04 · Results across five tasks

From Sokoban to 3D control, PIVOT sets a new bar

We evaluate on cognitive grids (Sokoban, FrozenLake), embodied 3D control (Navigation, PrimitiveSkill), and generative SVG. Full PIVOT is the M+P+R stack, and it ships with no extra test-time machinery.

+8% 3B over SFT-GRPO, 0.83 to 0.90
+5% 3B over GLANCE-Full, 0.86 to 0.90
+12% 2B over SFT-GRPO, 0.82 to 0.92
Method Cognitive Grid Embodied 3D Control Generative All
Sok. Frozen
Lake
Navigation PrimitiveSkill SVG
BaseCom.Avg. PlaceStackDraw.AlignAvg. DINODSAvg.
Open-source VLMs
Qwen2.5-VL-72B0.200.440.700.770.741.000.500.001.000.630.840.620.730.55
Qwen2.5-VL-7B0.140.140.330.380.350.000.000.000.750.190.840.270.560.28
Qwen2.5-VL-3B0.130.140.200.260.230.000.000.000.000.000.790.300.540.21
VLM-R1-3B0.160.150.330.340.340.000.000.000.000.000.790.270.540.24
Proprietary VLMs
o4-mini0.440.820.750.750.751.000.500.000.750.560.900.660.780.67
GPT-4o0.430.540.750.690.720.500.630.000.880.500.910.690.800.60
Gemini 2.5 Pro0.580.780.630.630.630.630.630.000.750.500.930.780.860.67
Claude 4.5 Sonnet0.310.800.670.670.670.630.500.001.000.530.950.810.880.64
Claude 3.7 Sonnet0.250.690.480.470.470.630.130.001.000.440.940.770.850.54
Previous SOTA / OPD baseline
VAGEN-Full0.790.720.800.810.811.000.881.001.000.970.900.660.780.81
GLANCE-Full0.850.780.860.880.871.000.881.001.000.970.920.700.810.86
SDAR0.520.770.880.840.861.001.001.001.001.000.910.680.800.79
PIVOT (Qwen2.5-VL-3B-Instruct)
Vanilla-GRPO0.780.710.870.750.810.920.860.890.930.900.900.660.780.80
SFT-GRPO0.820.750.930.770.850.960.900.930.970.940.920.680.800.83
M0.860.830.940.780.861.001.001.001.001.000.930.730.830.88
M+P0.900.770.950.830.891.001.001.001.001.000.930.740.840.88
M+P+R0.95+16%0.84+12%0.97+4%0.83+8%0.90+6%1.00+4%1.00+11%1.00+8%1.00+3%1.00+6%0.93+1%0.71+4%0.82+3%0.90+8%

Cyan deltas are gains of M+P+R over SFT-GRPO. Enriching Teacher context from M to M+P to M+P+R lifts overall accuracy to 0.90 (+8% over SFT-GRPO, +5% over GLANCE-Full).

Ablation: localization is not optional

A panel alone or a failure mode alone already helps. Combining them raises Sokoban / Navigation to about 0.90 / 0.89, and full M+P+R peaks at 0.95 / 0.90.

Replace the localized step with a random index (M+P-Random+R) and scores fall to 0.75 / 0.80, below Vanilla-GRPO. Coarse context is noise. Precise pivot localization is the ingredient that matters.

Pivot-step ablation
Variant Sok. Nav.
Vanilla-GRPO 0.78 0.81
P 0.86 0.85
M 0.86 0.86
M+P 0.90 0.89
M+P+R 0.95 0.90
M+P-Random+R 0.75 0.80

Training gets cleaner, not slower

Self-evolving pivot localization

On FrozenLake we can check Analyzer accuracy against a shortest-path oracle. Under M+P+R it climbs from about 0.56 toward 0.95 as training proceeds. Better localization feeds better credit, without a separate diagnostic model.

FrozenLake pivot accuracy rising from about 0.56 toward 0.95 over RL training.
FrozenLake t̂ accuracy. Gray is each step; green is a 15-step average.

Training efficiency

PIVOT kills bad actions earlier, so failed rollouts shorten. On Sokoban each RL step takes 95s versus 102s for Vanilla-GRPO, and far below a frozen 8B teacher at 373s. At test time the diagnostic branches are gone, so there is no parameter or latency tax.

Bar chart of seconds per RL step: vanilla GRPO 102s, PIVOT 95s, frozen 8B teacher 373s.
Time per RL step (s) on Sokoban.

05 · Findings, solution, outlook

What stalls GRPO, what PIVOT fixes, and where this goes next

Key findings

Where GRPO stalls, and what hindsight really buys

GRPO stalls in multi-turn visual tasks when a single-step error causes full-trajectory failure and gradient collapse. Hindsight distillation such as OPD/OPSD relies mainly on physical environment rollbacks at the exact pivot turn t*, while text hints add little.

Proposed solution · PIVOT

Bake hindsight into the weights

To break the rollback paradox, PIVOT puts hindsight into model weights rather than inference prompts. One VLM acts as Analyzer (locating the pivot), Teacher (distilling visual feedback), and Student in training, then deploys as a standalone Student with zero test-time overhead, lifting Qwen2.5-VL-3B to 0.90 (+8%).

Future outlook

Unsupervised pivot sensing

Replacing heuristic pivot labeling with unsupervised sensing via token entropy or self-play can drop simulator-rewind dependencies and open scalable post-training for open-world agents.

BibTeX

Cite this work

@misc{zhou2026pivotpivotawarepolicyself,
      title={PIVOT: Pivot-Aware On Policy Self Distillation for Multi-Turn VLM Agents}, 
      author={Jiazhou Zhou and Hu Zhou and Yucheng Chen and Jinyuan Qu and Ying-Cong Chen and Lei Zhang},
      year={2026},
      eprint={2609.35303},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.35303}, 
}