gaussian-policy and ppo: CleanRL's continuous-action PPO - #85
Merged
Merged
Conversation
added 4 commits
September 27, 2026 17:35
The coverage figure pairs expressive policies with algorithms built for
them and has no row for the baseline they are measured against. This adds
it, written to CleanRL's ppo_continuous_action.py with every default:
gaussian-policy tanh MLPs of 64 for the mean and the value, a log std
independent of the observation, orthogonal init (gain
sqrt 2; 0.01 on the mean's head, 1.0 on the value's).
ppo 2048-step rollouts, 10 epochs of 64-sample minibatches,
clip 0.2 on the ratio and the value, per-minibatch
advantage normalisation, one Adam at eps 1e-5, gradient
norm 0.5, lr 3e-4 annealed linearly over 488 iterations.
CleanRL clips actions and normalises observations and rewards in gymnasium
wrappers, and PlugRL's client wraps nothing, so they happen on the server:
the env receives the clipped sample and the buffer keeps the drawn one;
observation statistics are buffers updated after each learn, so collection
and learning share one normalisation; PPOBuffer divides rewards by the
running deviation of the discounted return and clips at 10, restarting the
return with each episode.
tests/test_gaussian_ppo.py checks each piece and that the pair learns the
bandit FPO's test uses and keeps it on five seeds (last/best at most 1.12).
Both servers call feedback for step t with terminated/truncated set to step t-1's outcome and next_terminated/next_truncated to step t's, and chain an episode's first frame to the last frame of the episode before. A frame carrying `terminated` therefore begins an episode. Until 48042e5 the GAE recursion cut a linked frame with the successor's dones[next_idx] - step t's own outcome. 48042e5 made it read the frame's own dones[step] (terminated[step]/truncated[step]) - step t-1's - so since then the step that ended an episode bootstrapped from the next episode's first state and kept accumulating its advantages, and each episode's first step was cut off from the rest. It also made a chain end carrying `dones` count as terminal and drop its bootstrap value. Both branches now read whether a frame's own transition ended from next_*. Three 3-step episodes of reward 1 at zero value, gamma = lambda = 1: main gave 4 3 2 1 3 2 1 2 1; the answer is 3 2 1 three times.
…ation; ppo refuses it
This was referenced Sep 27, 2026
Merged
added 2 commits
September 27, 2026 18:48
…e bit The factors are powers of two, but Adam's eps is added unscaled and leaves a trace in the last bits of elements whose gradient is within a few orders of it. After this branch changed the toy's returns, one parameter came out a few bits apart on CI and identical on Windows and on a Linux workstation, all torch 2.7.1.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #86 (the GAE episode-boundary fix), whose commit this branch contains: merge #86 first.
The coverage figure's rows pair expressive policies (flow, diffusion, pi0.5) with algorithms built for them. It has no row for the baseline those are measured against: a Gaussian MLP trained by PPO. This adds the pair, written to CleanRL's
ppo_continuous_action.pywith every default, so that if it fails on MuJoCo the platform failed, not an unfamiliar algorithm.gaussian-policy: tanh MLPs of 64×64 for the mean and the value, and a log std that does not depend on the observation. Weights are initialised orthogonally at gain √2, with 0.01 on the mean's last layer and 1.0 on the value's.--policy.deterministicacts with the mean, for evaluation, since the coverage clips act without training's sampling noise;pporefuses a deterministic policy.ppo: rollouts of 2048 steps, 10 epochs of 64-sample minibatches, and a 0.2 clip on both the ratio and the value. Advantages are normalised per minibatch. One Adam (eps 1e-5) covers the actor and critic, the gradient norm is clipped at 0.5, and the learning rate of 3e-4 anneals linearly to zero over 488 iterations (one million steps).CleanRL clips actions and normalises observations and rewards in gymnasium wrappers. PlugRL's env client wraps nothing, so all of that happens on the server:
PPOBufferdivides rewards by the running deviation of the discounted return and clips at ±10. The discounted return restarts with each episode, as in NormalizeReward and DPPO's RunningRewardScaler.DPPOBuffer's return carries across episodes because it follows the server's chain of frames. This PR leaves that alone.tests/test_gaussian_ppo.py(26 tests) covers:It also runs the pair on the bandit that
test_fpo_learns_anythinggives FPO, on five seeds. The pair goes from about -0.78 to between -0.120 and -0.160, and final over best is 1.00–1.14. FPO's hold test fails some of its ten cases on the same bar. That test uses lr 3e-3: at 3e-4, thirty small iterations are too few for the deviation to narrow.The full suite passes locally (261 passed, 3 skipped, before #86 was merged in; the PPO tests pass after it too). The MuJoCo run is next: E38, pre-registered, on HalfCheetah, Hopper and Walker2d with the figure's bars.