Skip to content

gaussian-policy and ppo: CleanRL's continuous-action PPO - #85

Merged
tactino merged 7 commits into
mainfrom
feat/gaussian-policy-ppo
Sep 27, 2026
Merged

tactino merged 7 commits into
mainfrom
feat/gaussian-policy-ppo

Conversation

@tactino

@tactino tactino commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Stacked on #86 (the GAE episode-boundary fix), whose commit this branch contains: merge #86 first.

The coverage figure's rows pair expressive policies (flow, diffusion, pi0.5) with algorithms built for them. It has no row for the baseline those are measured against: a Gaussian MLP trained by PPO. This adds the pair, written to CleanRL's ppo_continuous_action.py with every default, so that if it fails on MuJoCo the platform failed, not an unfamiliar algorithm.

  • gaussian-policy: tanh MLPs of 64×64 for the mean and the value, and a log std that does not depend on the observation. Weights are initialised orthogonally at gain √2, with 0.01 on the mean's last layer and 1.0 on the value's. --policy.deterministic acts with the mean, for evaluation, since the coverage clips act without training's sampling noise; ppo refuses a deterministic policy.
  • ppo: rollouts of 2048 steps, 10 epochs of 64-sample minibatches, and a 0.2 clip on both the ratio and the value. Advantages are normalised per minibatch. One Adam (eps 1e-5) covers the actor and critic, the gradient norm is clipped at 0.5, and the learning rate of 3e-4 anneals linearly to zero over 488 iterations (one million steps).

CleanRL clips actions and normalises observations and rewards in gymnasium wrappers. PlugRL's env client wraps nothing, so all of that happens on the server:

  • The environment gets the clipped sample. The buffer keeps the unclipped sample whose density the ratio is taken of.
  • Observation statistics are policy buffers, updated after each learn the way DPPO updates fpo-policy's. One iteration therefore collects and learns under a single normalisation, and the ratio is exactly 1 before the first update. CleanRL updates its statistics every step; ours lag by one rollout.
  • PPOBuffer divides rewards by the running deviation of the discounted return and clips at ±10. The discounted return restarts with each episode, as in NormalizeReward and DPPO's RunningRewardScaler. DPPOBuffer's return carries across episodes because it follows the server's chain of frames. This PR leaves that alone.

tests/test_gaussian_ppo.py (26 tests) covers:

  • CleanRL's initialisation
  • clipped action versus stored sample
  • recomputed log-probability equal to the stored one
  • observation normalisation and the ±10 clip
  • the return restarting per episode, and reward scaling
  • statistics updated only after learning
  • the annealing schedule
  • one Adam with gradient clip 0.5
  • checkpoint resume, including the reward scale and the schedule's position
  • the deterministic flag, and PPO refusing it
  • CLI registration in a fresh interpreter

It also runs the pair on the bandit that test_fpo_learns_anything gives FPO, on five seeds. The pair goes from about -0.78 to between -0.120 and -0.160, and final over best is 1.00–1.14. FPO's hold test fails some of its ten cases on the same bar. That test uses lr 3e-3: at 3e-4, thirty small iterations are too few for the deviation to narrow.

The full suite passes locally (261 passed, 3 skipped, before #86 was merged in; the PPO tests pass after it too). The MuJoCo run is next: E38, pre-registered, on HalfCheetah, Hopper and Walker2d with the figure's bars.

tactino added 4 commits September 27, 2026 17:35
The coverage figure pairs expressive policies with algorithms built for
them and has no row for the baseline they are measured against. This adds
it, written to CleanRL's ppo_continuous_action.py with every default:

  gaussian-policy  tanh MLPs of 64 for the mean and the value, a log std
                   independent of the observation, orthogonal init (gain
                   sqrt 2; 0.01 on the mean's head, 1.0 on the value's).
  ppo              2048-step rollouts, 10 epochs of 64-sample minibatches,
                   clip 0.2 on the ratio and the value, per-minibatch
                   advantage normalisation, one Adam at eps 1e-5, gradient
                   norm 0.5, lr 3e-4 annealed linearly over 488 iterations.

CleanRL clips actions and normalises observations and rewards in gymnasium
wrappers, and PlugRL's client wraps nothing, so they happen on the server:
the env receives the clipped sample and the buffer keeps the drawn one;
observation statistics are buffers updated after each learn, so collection
and learning share one normalisation; PPOBuffer divides rewards by the
running deviation of the discounted return and clips at 10, restarting the
return with each episode.

tests/test_gaussian_ppo.py checks each piece and that the pair learns the
bandit FPO's test uses and keeps it on five seeds (last/best at most 1.12).
Both servers call feedback for step t with terminated/truncated set to step
t-1's outcome and next_terminated/next_truncated to step t's, and chain an
episode's first frame to the last frame of the episode before. A frame
carrying `terminated` therefore begins an episode.

Until 48042e5 the GAE recursion cut a linked frame with the successor's
dones[next_idx] - step t's own outcome. 48042e5 made it read the frame's own
dones[step] (terminated[step]/truncated[step]) - step t-1's - so since then
the step that ended an episode bootstrapped from the next episode's first
state and kept accumulating its advantages, and each episode's first step
was cut off from the rest. It also made a chain end carrying `dones` count
as terminal and drop its bootstrap value.

Both branches now read whether a frame's own transition ended from next_*.
Three 3-step episodes of reward 1 at zero value, gamma = lambda = 1: main
gave 4 3 2 1 3 2 1 2 1; the answer is 3 2 1 three times.
tactino added 2 commits September 27, 2026 18:48
…e bit

The factors are powers of two, but Adam's eps is added unscaled and leaves a
trace in the last bits of elements whose gradient is within a few orders of
it. After this branch changed the toy's returns, one parameter came out a
few bits apart on CI and identical on Windows and on a Linux workstation,
all torch 2.7.1.
@tactino
tactino merged commit 57c9bd3 into main Sep 27, 2026
3 checks passed
@tactino
tactino deleted the feat/gaussian-policy-ppo branch September 27, 2026 22:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant