Skip to content

dppo-gaussian-policy and ppo dppo-square: DPPO's Gaussian PPO on square - #92

Merged
tactino merged 1 commit into
mainfrom
feat/dppo-gaussian-square
Sep 28, 2026
Merged

tactino merged 1 commit into
mainfrom
feat/dppo-gaussian-square

Conversation

@tactino

@tactino tactino commented Sep 28, 2026

Copy link
Copy Markdown
Member

The coverage figure's Gaussian · PPO row (#85, E38) has no square cell. From random weights, square's sparse reward gives a Gaussian nothing to learn from. DPPO's own baseline instead fine-tunes a Gaussian MLP with PPO from a pretrained checkpoint it releases. This PR adds that policy and that setting, so the cell can be run the way E34 ran DPPO's diffusion policy.

dppo-gaussian-policy is DPPO's Gaussian_MLP with a fixed, learned std, and GaussianModel's sampling. It is built on the dppo package's own ResidualMLP and CriticObs, so the released square checkpoint loads by its own keys. Loaded, it has 2,152,485 parameters, the paper's "2.15M".

  • Mean: tanh over a residual Mish MLP, [23, 1024, 1024, 1024, 28]. Std: one learned log-variance per action dimension, clamped, and repeated over the chunk of 4.
  • Samples are clamped to mean ± 3σ.
  • The logprob is the mean over the chunk's 28 elements, clamped to [−5, 2], as PPO_Gaussian scores it.
  • Observations and actions go through DPPO's normalization.npz, as dppo-policy does.
  • Loads the checkpoint's model weights, not ema, as GaussianModel does. The only key it may miss is network.logvar; any other missing key is refused.
  • Kept on purpose: the released checkpoint's logvar_max is 0 (σ ≤ 1), from pretraining's default. DPPO's non-strict load puts it over the fine-tuning config's 0.2, so DPPO's fine-tuning actually bounds σ at 1.0. This policy does the same, and logs the bound.

ppo gains four options, all off by default:

  • critic_learning_rate: a second parameter group; annealing scales both.
  • n_critic_warmup_itrs: the actor gets no gradient.
  • max_grad_norm: float | None: None means no clipping.
  • adam_eps and reward_scaling_gamma: the running return's own discount.

ppo dppo-square carries cfg/robomimic/finetune/square/ft_ppo_gaussian_mlp.yaml at irom-lab/dppo cc7234ad:

  • actor learning rate 1e-4, critic 1e-3, both constant
  • 20,000 chunks per iteration, minibatches of 10,000, 10 epochs
  • γ 0.999, λ 0.95, clip 0.01, target KL 1, value coefficient 0.5
  • no value clip, no gradient clip
  • reward scaling by a 0.99 running return, clipped at 10
  • one critic-only iteration

The paper's table lists an actor rate of 1e-5; the config used that before v0.7. Not carried over:

  • DPPO's evaluation-only iterations, which train nothing.
  • DPPO's GAE masks only termination, so it bootstraps across the time-outs that end every square episode. Here a time-out ends the episode for GAE, as for every other algorithm.

Tests:

  • tests/test_dppo_gaussian_policy.py (13): keys and shapes; loading, including logvar_max, model over ema, and refusing a checkpoint that is missing a layer; the action chunk in the environment's units; min-max observation scaling; the 3σ clamp; the clamped mean logprob; learning scoring what was sampled; deterministic mode.
  • tests/test_ppo_dppo_square.py (10): the variant's values; the split optimizer and its annealing; the warmup; no clipping; the reward-scaling discount; and PPO driving dppo-gaussian-policy end to end.

Full suite: 293 passed, 3 skipped. Both files skip where the dppo extra is missing.

E40, which runs this on square, is next (exp/e40-square-gaussian-ppo, stacked on this).

DPPO fine-tunes a Gaussian MLP with PPO as its robomimic baseline, from
pretrained checkpoints it releases. This adds that policy, built on the dppo
package's ResidualMLP and CriticObs so the released square checkpoint loads
by its own keys (2,152,485 parameters, as the paper states), and the options
ppo needs to run DPPO's square config:

  critic_learning_rate   a second parameter group (DPPO: actor 1e-4, critic 1e-3)
  n_critic_warmup_itrs   iterations in which the actor gets no gradient
  max_grad_norm = None   no clipping
  adam_eps               CleanRL's 1e-5 by default; DPPO keeps 1e-8
  reward_scaling_gamma   the running return's own discount (DPPO: 0.99)

`ppo dppo-square` carries ft_ppo_gaussian_mlp.yaml's values at irom-lab/dppo
cc7234ad. The released checkpoint's logvar_max (a deviation of 1) replaces the
config's 0.2 on load, as in DPPO; the policy keeps that and says so.
@tactino
tactino merged commit 01dd73d into main Sep 28, 2026
3 checks passed
@tactino
tactino deleted the feat/dppo-gaussian-square branch September 28, 2026 19:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant