Skip to content

figure: the gaussian-policy · PPO row (E38) - #88

Merged
tactino merged 16 commits into
mainfrom
figure/gaussian-ppo
Sep 27, 2026
Merged

tactino merged 16 commits into
mainfrom
figure/gaussian-ppo

Conversation

@tactino

@tactino tactino commented Sep 27, 2026

Copy link
Copy Markdown
Member

Stacked on #87 (E38), which sits on #85 and #86. Merge those first; the site PR goes last.

This adds the fourth row to the coverage figure: gaussian-policy · PPO, learns on HalfCheetah, Hopper and Walker2d (E38).

  • cells.json: the row and its three cells. Each cell's seed is the median of the last ten iterations: 0, 1 and 2. The clips come from the final checkpoints, evaluated with --policy.deterministic so that, like the other rows, they act without training's sampling noise. Median evaluation returns are 1681, 3063 and 3907. The row has no square cell because it was not run there; the page leaves that slot empty, as it already could.
  • commands.json: the three training commands. The source is experiments/e38-gaussian-ppo/run.sh.
  • grid_still.py: the label column now widens to the longest row label (at least its old width). "gaussian-policy" ran into the first cell.
  • README.md: the image's alt text was out of date even before this row: it still said dppo-policy was "still rising" and fpo-policy under DPPO "has not learned". It now describes the figure as it stands. The Hopper command comparison gains the fourth server line, which the README command test parses.

Built on guangzhao in the fig worktree at 65eecb9 with run.sh for the three cells, then data.py and grid_still.py.

tactino added 16 commits September 27, 2026 17:35
The coverage figure pairs expressive policies with algorithms built for
them and has no row for the baseline they are measured against. This adds
it, written to CleanRL's ppo_continuous_action.py with every default:

  gaussian-policy  tanh MLPs of 64 for the mean and the value, a log std
                   independent of the observation, orthogonal init (gain
                   sqrt 2; 0.01 on the mean's head, 1.0 on the value's).
  ppo              2048-step rollouts, 10 epochs of 64-sample minibatches,
                   clip 0.2 on the ratio and the value, per-minibatch
                   advantage normalisation, one Adam at eps 1e-5, gradient
                   norm 0.5, lr 3e-4 annealed linearly over 488 iterations.

CleanRL clips actions and normalises observations and rewards in gymnasium
wrappers, and PlugRL's client wraps nothing, so they happen on the server:
the env receives the clipped sample and the buffer keeps the drawn one;
observation statistics are buffers updated after each learn, so collection
and learning share one normalisation; PPOBuffer divides rewards by the
running deviation of the discounted return and clips at 10, restarting the
return with each episode.

tests/test_gaussian_ppo.py checks each piece and that the pair learns the
bandit FPO's test uses and keeps it on five seeds (last/best at most 1.12).
Both servers call feedback for step t with terminated/truncated set to step
t-1's outcome and next_terminated/next_truncated to step t's, and chain an
episode's first frame to the last frame of the episode before. A frame
carrying `terminated` therefore begins an episode.

Until 48042e5 the GAE recursion cut a linked frame with the successor's
dones[next_idx] - step t's own outcome. 48042e5 made it read the frame's own
dones[step] (terminated[step]/truncated[step]) - step t-1's - so since then
the step that ended an episode bootstrapped from the next episode's first
state and kept accumulating its advantages, and each episode's first step
was cut off from the rest. It also made a chain end carrying `dones` count
as terminal and drop its bootstrap value.

Both branches now read whether a frame's own transition ended from next_*.
Three 3-step episodes of reward 1 at zero value, gamma = lambda = 1: main
gave 4 3 2 1 3 2 1 2 1; the answer is 3 2 1 three times.
…er and Walker2d, 3 of 3 each, at CleanRL's returns
…e bit

The factors are powers of two, but Adam's eps is added unscaled and leaves a
trace in the last bits of elements whose gradient is within a few orders of
it. After this branch changed the toy's returns, one parameter came out a
few bits apart on CI and identical on Windows and on a Linux workstation,
all torch 2.7.1.
@tactino
tactino merged commit efdab35 into main Sep 27, 2026
3 checks passed
@tactino
tactino deleted the figure/gaussian-ppo branch September 27, 2026 22:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant