Skip to content

E38: gaussian-policy under ppo learns HalfCheetah, Hopper and Walker2d at CleanRL's returns - #87

Merged
tactino merged 13 commits into
mainfrom
exp/e38-gaussian-ppo
Sep 27, 2026
Merged

tactino merged 13 commits into
mainfrom
exp/e38-gaussian-ppo

Conversation

@tactino

@tactino tactino commented Sep 27, 2026

Copy link
Copy Markdown
Member

Stacked on #85 (gaussian-policy and ppo), which contains #86 (the GAE fix). Merge #86, then #85, then this.

E38 was pre-registered at 2f70733, after a three-iteration pilot and the GAE diagnostic. It runs gaussian-policy under ppo with every CleanRL default on HalfCheetah, Hopper and Walker2d: three seeds, 488 iterations of 2,048 steps, on guangzhao.

task iterations 479-488, seeds 0 / 1 / 2 bar CleanRL (-v4, 1M steps)
HalfCheetah +1848 / +1827 / +1952 over the first (1443–1558) +200 1442.64 ± 46.03
Hopper 2299 / 2207 / 2181 500 2382.86 ± 271.74
Walker2d 2957 / 3260 / 3154 500 2287.95 ± 571.78
  • P1 (runs end to end) holds on all nine runs. The three cells took 28 minutes, run together.
  • V1 passes: every iteration learned at CleanRL's annealed rate.
  • P2–P4 hold. It learns all three tasks, 3 of 3 seeds each, and the smallest figure is 4.4 times its bar.

The row gaussian-policy · PPO learns on all three MuJoCo tasks. The returns sit about where CleanRL's do. That was reported, not tested: the tasks are -v5, and three differences were declared in advance. Walker2d ends above CleanRL's spread, and E38 does not explain that. Square was not run.

Checkpoints and tensorboards stay on guangzhao. The logs, verdicts.txt and summary.tsv are here. Adding the row to the coverage figure comes next, in its own PR.

tactino added 10 commits September 27, 2026 17:35
The coverage figure pairs expressive policies with algorithms built for
them and has no row for the baseline they are measured against. This adds
it, written to CleanRL's ppo_continuous_action.py with every default:

  gaussian-policy  tanh MLPs of 64 for the mean and the value, a log std
                   independent of the observation, orthogonal init (gain
                   sqrt 2; 0.01 on the mean's head, 1.0 on the value's).
  ppo              2048-step rollouts, 10 epochs of 64-sample minibatches,
                   clip 0.2 on the ratio and the value, per-minibatch
                   advantage normalisation, one Adam at eps 1e-5, gradient
                   norm 0.5, lr 3e-4 annealed linearly over 488 iterations.

CleanRL clips actions and normalises observations and rewards in gymnasium
wrappers, and PlugRL's client wraps nothing, so they happen on the server:
the env receives the clipped sample and the buffer keeps the drawn one;
observation statistics are buffers updated after each learn, so collection
and learning share one normalisation; PPOBuffer divides rewards by the
running deviation of the discounted return and clips at 10, restarting the
return with each episode.

tests/test_gaussian_ppo.py checks each piece and that the pair learns the
bandit FPO's test uses and keeps it on five seeds (last/best at most 1.12).
Both servers call feedback for step t with terminated/truncated set to step
t-1's outcome and next_terminated/next_truncated to step t's, and chain an
episode's first frame to the last frame of the episode before. A frame
carrying `terminated` therefore begins an episode.

Until 48042e5 the GAE recursion cut a linked frame with the successor's
dones[next_idx] - step t's own outcome. 48042e5 made it read the frame's own
dones[step] (terminated[step]/truncated[step]) - step t-1's - so since then
the step that ended an episode bootstrapped from the next episode's first
state and kept accumulating its advantages, and each episode's first step
was cut off from the rest. It also made a chain end carrying `dones` count
as terminal and drop its bootstrap value.

Both branches now read whether a frame's own transition ended from next_*.
Three 3-step episodes of reward 1 at zero value, gamma = lambda = 1: main
gave 4 3 2 1 3 2 1 2 1; the answer is 3 2 1 three times.
…er and Walker2d, 3 of 3 each, at CleanRL's returns
tactino added 3 commits September 27, 2026 18:48
…e bit

The factors are powers of two, but Adam's eps is added unscaled and leaves a
trace in the last bits of elements whose gradient is within a few orders of
it. After this branch changed the toy's returns, one parameter came out a
few bits apart on CI and identical on Windows and on a Linux workstation,
all torch 2.7.1.
@tactino
tactino merged commit 065e980 into main Sep 27, 2026
3 checks passed
@tactino
tactino deleted the exp/e38-gaussian-ppo branch September 27, 2026 22:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant