dppo-gaussian-policy and ppo dppo-square: DPPO's Gaussian PPO on square - #92
Merged
Merged
Conversation
DPPO fine-tunes a Gaussian MLP with PPO as its robomimic baseline, from pretrained checkpoints it releases. This adds that policy, built on the dppo package's ResidualMLP and CriticObs so the released square checkpoint loads by its own keys (2,152,485 parameters, as the paper states), and the options ppo needs to run DPPO's square config: critic_learning_rate a second parameter group (DPPO: actor 1e-4, critic 1e-3) n_critic_warmup_itrs iterations in which the actor gets no gradient max_grad_norm = None no clipping adam_eps CleanRL's 1e-5 by default; DPPO keeps 1e-8 reward_scaling_gamma the running return's own discount (DPPO: 0.99) `ppo dppo-square` carries ft_ppo_gaussian_mlp.yaml's values at irom-lab/dppo cc7234ad. The released checkpoint's logvar_max (a deviation of 1) replaces the config's 0.2 on load, as in DPPO; the policy keeps that and says so.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The coverage figure's Gaussian · PPO row (#85, E38) has no square cell. From random weights, square's sparse reward gives a Gaussian nothing to learn from. DPPO's own baseline instead fine-tunes a Gaussian MLP with PPO from a pretrained checkpoint it releases. This PR adds that policy and that setting, so the cell can be run the way E34 ran DPPO's diffusion policy.
dppo-gaussian-policyis DPPO'sGaussian_MLPwith a fixed, learned std, andGaussianModel's sampling. It is built on thedppopackage's ownResidualMLPandCriticObs, so the released square checkpoint loads by its own keys. Loaded, it has 2,152,485 parameters, the paper's "2.15M".tanhover a residual Mish MLP, [23, 1024, 1024, 1024, 28]. Std: one learned log-variance per action dimension, clamped, and repeated over the chunk of 4.PPO_Gaussianscores it.normalization.npz, asdppo-policydoes.modelweights, notema, asGaussianModeldoes. The only key it may miss isnetwork.logvar; any other missing key is refused.logvar_maxis 0 (σ ≤ 1), from pretraining's default. DPPO's non-strict load puts it over the fine-tuning config's 0.2, so DPPO's fine-tuning actually bounds σ at 1.0. This policy does the same, and logs the bound.ppogains four options, all off by default:critic_learning_rate: a second parameter group; annealing scales both.n_critic_warmup_itrs: the actor gets no gradient.max_grad_norm: float | None:Nonemeans no clipping.adam_epsandreward_scaling_gamma: the running return's own discount.ppo dppo-squarecarriescfg/robomimic/finetune/square/ft_ppo_gaussian_mlp.yamlat irom-lab/dppo cc7234ad:The paper's table lists an actor rate of 1e-5; the config used that before v0.7. Not carried over:
Tests:
tests/test_dppo_gaussian_policy.py(13): keys and shapes; loading, includinglogvar_max,modeloverema, and refusing a checkpoint that is missing a layer; the action chunk in the environment's units; min-max observation scaling; the 3σ clamp; the clamped mean logprob; learning scoring what was sampled; deterministic mode.tests/test_ppo_dppo_square.py(10): the variant's values; the split optimizer and its annealing; the warmup; no clipping; the reward-scaling discount; and PPO drivingdppo-gaussian-policyend to end.Full suite: 293 passed, 3 skipped. Both files skip where the
dppoextra is missing.E40, which runs this on square, is next (
exp/e40-square-gaussian-ppo, stacked on this).