Skip to content

FPO++'s optimizer, gradient clipping, advantage normalisation and Huber CFM loss, as options - #91

Merged
tactino merged 1 commit into
mainfrom
feat/fpo-plus-plus-optimizer
Sep 28, 2026
Merged

tactino merged 1 commit into
mainfrom
feat/fpo-plus-plus-optimizer

Conversation

@tactino

@tactino tactino commented Sep 28, 2026

Copy link
Copy Markdown
Member

E37 (#89) fine-tuned a behaviour-cloned fpo-policy on square with FPO++'s policy loss. Over fifty evaluation episodes it fell from 0.50 to 0.24–0.42, while DPPO took the same clone to 0.78–0.90. The randomly initialised critic never fit the returns. That run had only FPO++'s loss. FPO++'s square fine-tuning (amazon-far/fpo-control, manipulation_experiments/finetune_online_rl.py, read for facts, nothing copied) also has the following. Each is now an option, off by default:

option FPO++ on square off (FPO until now)
critic_learning_rate, adam_eps, weight_decay, actor_adam_beta2 two AdamW groups: actor 1e-5, betas (0.9, 0.99); critic 1e-4; eps 1e-5; weight decay 1e-6 one Adam over both
max_grad_norm actor and critic clipped separately, 25 no clipping
normalize_advantage_per_minibatch per minibatch (375 chunks) over the buffer
cfm_loss_huber_delta Huber, δ = 1: d² inside, 2δ|d| − δ² beyond squared error

Two details:

  • During a critic warmup, the actor's AdamW group gets no gradient at all, not a zero one. AdamW decays every parameter it steps, so a zero gradient would still move the actor.
  • FPO's learn step now reports the critic's explained variance under rollout/, as DPPO always has. E37 could not read it.

tests/test_fpo_plus_plus_optimizer.py has 17 tests:

  • the groups, their rates, betas, eps and decay, and that the default is still the one Adam
  • the actor bit-identical through a warmup with weight decay on, and moving after it
  • checkpoint round-trip
  • clipping called per group with the right parameters and limit, and a tight clip bounding the step
  • per-minibatch normalisation giving unit spread per minibatch, against a control where the buffer-level one does not
  • the Huber formula
  • explained variance with and without fpo_playground_trick

The one source-inspection test that forbids normalising inside _compute_loss still holds by default. The opt-in goes through _minibatch_advantage, and that test's docstring now says so. Full suite: 287 passed, 3 skipped.

E39, which runs these on square, is next (exp/e39-square-fpo-plus-plus, stacked on this).

…er CFM loss, as options

E37 ran FPO++'s policy loss on square from a behaviour-cloned start and the
policy fell from 0.50 to 0.24-0.42 over fifty evaluation episodes, while its
randomly initialised critic never fit the returns. FPO++'s square fine-tuning
(amazon-far/fpo-control, manipulation_experiments/finetune_online_rl.py) has
more than the loss, and this adds the rest, each off by default:

  critic_learning_rate, adam_eps, weight_decay, actor_adam_beta2
      two AdamW groups (FPO++: actor 1e-5 betas (0.9, 0.99), critic 1e-4,
      eps 1e-5, weight decay 1e-6); left at their defaults, FPO's one Adam.
      During a critic warmup the actor's group gets no gradient at all, so
      weight decay cannot move it.
  max_grad_norm
      the actor's and the critic's gradients clipped separately (FPO++: 25);
      their pre-clip maxima are logged.
  normalize_advantage_per_minibatch
      as FPO++ does, instead of over the buffer.
  cfm_loss_huber_delta
      FPO++'s Huber on the flow-matching error, d^2 within delta and
      2 delta |d| - delta^2 beyond (FPO++: 1).

FPO's learn step now also reports the critic's explained variance, which
DPPO always has and E37 could not read.
@tactino
tactino merged commit 23b2ac1 into main Sep 28, 2026
3 checks passed
@tactino
tactino deleted the feat/fpo-plus-plus-optimizer branch September 28, 2026 19:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant