figure: the gaussian-policy · PPO row (E38) - #88
Merged
Merged
Conversation
added 16 commits
September 27, 2026 17:35
The coverage figure pairs expressive policies with algorithms built for
them and has no row for the baseline they are measured against. This adds
it, written to CleanRL's ppo_continuous_action.py with every default:
gaussian-policy tanh MLPs of 64 for the mean and the value, a log std
independent of the observation, orthogonal init (gain
sqrt 2; 0.01 on the mean's head, 1.0 on the value's).
ppo 2048-step rollouts, 10 epochs of 64-sample minibatches,
clip 0.2 on the ratio and the value, per-minibatch
advantage normalisation, one Adam at eps 1e-5, gradient
norm 0.5, lr 3e-4 annealed linearly over 488 iterations.
CleanRL clips actions and normalises observations and rewards in gymnasium
wrappers, and PlugRL's client wraps nothing, so they happen on the server:
the env receives the clipped sample and the buffer keeps the drawn one;
observation statistics are buffers updated after each learn, so collection
and learning share one normalisation; PPOBuffer divides rewards by the
running deviation of the discounted return and clips at 10, restarting the
return with each episode.
tests/test_gaussian_ppo.py checks each piece and that the pair learns the
bandit FPO's test uses and keeps it on five seeds (last/best at most 1.12).
Both servers call feedback for step t with terminated/truncated set to step t-1's outcome and next_terminated/next_truncated to step t's, and chain an episode's first frame to the last frame of the episode before. A frame carrying `terminated` therefore begins an episode. Until 48042e5 the GAE recursion cut a linked frame with the successor's dones[next_idx] - step t's own outcome. 48042e5 made it read the frame's own dones[step] (terminated[step]/truncated[step]) - step t-1's - so since then the step that ended an episode bootstrapped from the next episode's first state and kept accumulating its advantages, and each episode's first step was cut off from the rest. It also made a chain end carrying `dones` count as terminal and drop its bootstrap value. Both branches now read whether a frame's own transition ended from next_*. Three 3-step episodes of reward 1 at zero value, gamma = lambda = 1: main gave 4 3 2 1 3 2 1 2 1; the answer is 3 2 1 three times.
…, on HalfCheetah, Hopper and Walker2d
…ation; ppo refuses it
…er and Walker2d, 3 of 3 each, at CleanRL's returns
…e bit The factors are powers of two, but Adam's eps is added unscaled and leaves a trace in the last bits of elements whose gradient is within a few orders of it. After this branch changed the toy's returns, one parameter came out a few bits apart on CI and identical on Windows and on a Linux workstation, all torch 2.7.1.
This was referenced Sep 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #87 (E38), which sits on #85 and #86. Merge those first; the site PR goes last.
This adds the fourth row to the coverage figure:
gaussian-policy· PPO, learns on HalfCheetah, Hopper and Walker2d (E38).cells.json: the row and its three cells. Each cell's seed is the median of the last ten iterations: 0, 1 and 2. The clips come from the final checkpoints, evaluated with--policy.deterministicso that, like the other rows, they act without training's sampling noise. Median evaluation returns are 1681, 3063 and 3907. The row has no square cell because it was not run there; the page leaves that slot empty, as it already could.commands.json: the three training commands. The source isexperiments/e38-gaussian-ppo/run.sh.grid_still.py: the label column now widens to the longest row label (at least its old width). "gaussian-policy" ran into the first cell.README.md: the image's alt text was out of date even before this row: it still said dppo-policy was "still rising" and fpo-policy under DPPO "has not learned". It now describes the figure as it stands. The Hopper command comparison gains the fourth server line, which the README command test parses.Built on
guangzhaoin thefigworktree at 65eecb9 withrun.shfor the three cells, thendata.pyandgrid_still.py.