Skip to content

E40: DPPO's Gaussian MLP learns square under DPPO's own Gaussian PPO - #94

Merged
tactino merged 4 commits into
mainfrom
exp/e40-square-gaussian-ppo
Sep 28, 2026
Merged

tactino merged 4 commits into
mainfrom
exp/e40-square-gaussian-ppo

Conversation

@tactino

@tactino tactino commented Sep 28, 2026

Copy link
Copy Markdown
Member

Stacked on #92 (dppo-gaussian-policy and ppo dppo-square), whose commit this branch contains: merge #92 first.

E40 was pre-registered at 456da16, after a two-iteration pilot. It runs DPPO's Gaussian MLP on robomimic square, started from DPPO's released checkpoint (its model weights), under every value of DPPO's ft_ppo_gaussian_mlp.yaml. Three seeds, 40 iterations of 80,000 steps, on guangzhao.

seed start (iterations 1-2) end (31-40) gain
0 0.185 0.547 +0.362
1 0.195 0.537 +0.342
2 0.212 0.565 +0.352
  • P1 (runs) and V1 (released model weights with σ ≤ 1, the dppo-square config) hold.
  • P2 holds: it learns, 3 of 3. It is still rising at 40.
  • The clip fraction sat near 0.45 throughout: a 0.01 clip on the chunk's mean log-probability binds that often in DPPO's own setting, and it learns anyway.

The Gaussian · PPO row's empty square cell becomes learns. That makes it the third pair to learn all four tasks. These are training-rollout rates, not DPPO's deterministic evaluations, and E40 does not claim to reproduce DPPO's curve.

DPPO fine-tunes a Gaussian MLP with PPO as its robomimic baseline, from
pretrained checkpoints it releases. This adds that policy, built on the dppo
package's ResidualMLP and CriticObs so the released square checkpoint loads
by its own keys (2,152,485 parameters, as the paper states), and the options
ppo needs to run DPPO's square config:

  critic_learning_rate   a second parameter group (DPPO: actor 1e-4, critic 1e-3)
  n_critic_warmup_itrs   iterations in which the actor gets no gradient
  max_grad_norm = None   no clipping
  adam_eps               CleanRL's 1e-5 by default; DPPO keeps 1e-8
  reward_scaling_gamma   the running return's own discount (DPPO: 0.99)

`ppo dppo-square` carries ft_ppo_gaussian_mlp.yaml's values at irom-lab/dppo
cc7234ad. The released checkpoint's logvar_max (a deviation of 1) replaces the
config's 0.2 on load, as in DPPO; the policy keeps that and says so.
@tactino
tactino merged commit bb4e5f4 into main Sep 28, 2026
3 checks passed
@tactino
tactino deleted the exp/e40-square-gaussian-ppo branch September 28, 2026 19:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant