Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 26 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -126,10 +126,15 @@ We use `uv` to manage dependencies and development environments.
- `dppo` - DPPO (Diffusion Policy Policy Optimization). No extras needed,
and it runs against `fpo-policy`.
- `dppo-dist` - Distributed DPPO (experimental)
- `ppo` - PPO as CleanRL's `ppo_continuous_action.py` runs it, every
default included. Drives `gaussian-policy`; no extras needed.
- `eval` - Evaluation only, no learning

**Policies:**
- `fpo-policy` - Flow policy. Defaults to `obs_dim=17`, `action_dim=6`
- `gaussian-policy` - CleanRL's Gaussian MLP: tanh layers of 64, a log std
that does not depend on the observation. Defaults to `obs_dim=17`,
`action_dim=6`
- `dummy-policy` - Outputs random actions (for testing)
- `dppo-policy` - DPPO policy (requires `plugrl-server[dppo]` and a checkpoint)
- `pi0-policy` - PI0 policy (OpenPI). Needs more than a checkpoint. The
Expand Down Expand Up @@ -213,6 +218,27 @@ million steps therefore learns exactly once, at the very end - producing a
single point rather than a curve. 4096 gives one update per 4096 environment
steps.

#### The baseline: a Gaussian policy with PPO

The pair every other one is measured against, written to CleanRL's
`ppo_continuous_action.py`: a rollout of 2048 steps, ten epochs of 32
minibatches, clip 0.2, a learning rate of 3e-4 annealed to zero over 488
iterations (a million steps), observations and rewards normalised. CleanRL
does the normalising and the action clipping in gymnasium wrappers; PlugRL's
client does not wrap, so the policy and the algorithm do it on the server.

```bash
# Terminal 1: Hopper-v5 has an 11-dimensional observation and 3 actions
python -m plugrl_server.cli gaussian-policy default ppo default \
--port 8000 --policy.device cpu \
--policy.obs-dim 11 --policy.action-dim 3

# Terminal 2
python -m plugrl_env_client.cli mujoco-v1 \
--server-port 8000 --num-envs 1 --num-episodes 100000 \
--env.name Hopper-v5 --runner.replan-steps 1 --runner.seed 0
```

#### Quick Start: Testing with Dummy Components

To test the agent-server connection with dummy algorithm and policy:
Expand Down
83 changes: 83 additions & 0 deletions experiments/e38-gaussian-ppo/FINDINGS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,83 @@
# E38: a Gaussian policy with PPO, run as CleanRL runs it, learns all three MuJoCo tasks through PlugRL - at CleanRL's returns

2026-09-27 · Linux workstation (`guangzhao`), CPU only · three cells, three
seeds, 488 iterations of 2,048 steps · protocol: [`PROTOCOL.md`](PROTOCOL.md)
(`2f70733`, after the pilot and the GAE diagnostic in `results/pilot.txt`,
before the registered run)

---

## The result

`gaussian-policy` under `ppo` (#85) is CleanRL's `ppo_continuous_action.py`
with every default, on code with the GAE episode-boundary fix (#86). The
coverage figure's rule, read on iterations 479-488:

| cell | seed | first iteration | iterations 479-488 | CleanRL (-v4) |
| --- | --- | --- | --- | --- |
| HalfCheetah | 0 | -335.2 | 1512.9 (**+1848**) | 1442.64 +/- 46.03 |
| | 1 | -384.6 | 1442.5 (**+1827**) | |
| | 2 | -393.9 | 1557.8 (**+1952**) | |
| Hopper | 0 | 12.1 | **2298.7** | 2382.86 +/- 271.74 |
| | 1 | 15.2 | **2207.4** | |
| | 2 | 14.2 | **2181.3** | |
| Walker2d | 0 | -0.7 | **2957.2** | 2287.95 +/- 571.78 |
| | 1 | -1.8 | **3260.0** | |
| | 2 | -0.6 | **3154.2** | |

* **P1 holds**: nine runs, each with 488 iterations logged, 24 checkpoints,
no traceback and a client that exited 0. The three cells ran at once
beside E37 and took 28 minutes.
* **V1 passes**: every iteration of every run learned at CleanRL's annealed
rate, from 3e-4 down to 6.15e-7 (3e-4 / 488) at the last.
* **P2, P3, P4 hold**: it learns Hopper, Walker2d and HalfCheetah, each on
3 of 3 seeds. The smallest figure is 4.4 times its bar.

The new row, `gaussian-policy` · PPO, **learns** on all three tasks.

---

## Beside CleanRL

Mean return over three windows:

| cell | seed | 91-100 | 241-250 | 479-488 |
| --- | --- | --- | --- | --- |
| HalfCheetah | 0 | 744.4 | 1322.5 | 1512.9 |
| | 1 | 999.0 | 1348.1 | 1442.5 |
| | 2 | 925.5 | 1428.5 | 1557.8 |
| Hopper | 0 | 1623.2 | 2242.8 | 2298.7 |
| | 1 | 1297.1 | 2486.9 | 2207.4 |
| | 2 | 920.3 | 2440.6 | 2181.3 |
| Walker2d | 0 | 514.9 | 2747.0 | 2957.2 |
| | 1 | 476.6 | 2037.8 | 3260.0 |
| | 2 | 652.6 | 3344.2 | 3154.2 |

HalfCheetah ends at CleanRL's number or a little above it, and Hopper
within its spread. Walker2d ends above it, by more than CleanRL's own
spread; E38 does not explain that. The protocol registered these as reported
figures, not a test. The tasks are -v5, not -v4, and three things are done
differently (declared there). What they do show is that the one pair here
whose behaviour is known well outside this project comes out of PlugRL
about where it comes out of CleanRL. Here the server has no simulator
installed, and the environment runs in a separate client with only the
protocol between them.

Episodes grew from 18-20 steps to 586-636 on Hopper and 686-792 on Walker2d.
HalfCheetah's are always 1,000.

---

## What E38 does not show

* **That the platform was right before #86.** E38 ran with the GAE fix. The
diagnostic in `results/pilot.txt` ran this pair on Hopper for 100
iterations on each side of the fix. By iteration 100 the two were
indistinguishable on three seeds (951 against 941 on average), so the
defect cost this task little. That diagnostic was not registered.
* **The square column.** `gaussian-policy` was not run on robomimic square.
From random weights, square's sparse reward gives it nothing to learn from
(E27), and no behaviour-cloned Gaussian start was built.

Checkpoints and tensorboards stay on `guangzhao`. The logs, `verdicts.txt`
and `summary.tsv` are here.
128 changes: 128 additions & 0 deletions experiments/e38-gaussian-ppo/PROTOCOL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,128 @@
# E38 measurement protocol (pre-registered)

**Written 2026-09-27, before any E38 run. The three-iteration pilot and the
GAE diagnostic below came first and are recorded in `results/pilot.txt`.**

This file must not be edited after the first registered data point. Anything
learned afterwards goes in `AMENDMENT.md`, dated.

---

## The question

The coverage figure's rows are expressive policies with the algorithms built
for them. It has no row for the baseline they are measured against. #85 adds
one: `gaussian-policy` under `ppo`, written to CleanRL's
`ppo_continuous_action.py` with every default.

**Run as CleanRL runs it, does it learn HalfCheetah, Hopper and Walker2d
through PlugRL?**

### What this cannot settle

* Whether PlugRL reproduces CleanRL's returns. The tasks are -v5, not -v4,
and three things happen differently (declared below). The returns are
reported beside CleanRL's, not tested against them.
* Anything about the other rows.

---

## Declared in advance: what was already known

1. **CleanRL's own results** for this file, from its documentation (MuJoCo
-v4; the benchmark command passes no `--total-timesteps`, so its default
of 1,000,000; three seeds): HalfCheetah 1442.64 +/- 46.03, Hopper
2382.86 +/- 271.74, Walker2d 2287.95 +/- 571.78.
2. **The pair on a bandit** (#85's tests, with #86 merged): from about
-0.78 to between -0.120 and -0.160 on five seeds, final over best
1.00-1.14.
3. **The pilot**: three iterations on each task, seed 0, at ab6e32f. Every
run ended with exit 0 and no traceback, at 528-584 environment steps a
second including learning - about 30 minutes for 488 iterations alone.
4. **GAE ended every episode one step late** on every branch since 48042e5
(#86); E38 runs with the fix. A diagnostic of this pair on Hopper-v5, 100
iterations (annealed over 100), seeds 0-2, at ab6e32f and 274929c:
iterations 41-50 at 267 / 589 / 367 before the fix and 674 / 673 / 454
after; iterations 91-100 at 432 / 1512 / 910 before and 1005 / 868 / 949
after.
5. **Runs are not reproducible bit for bit** after the first update (E20).

### Where this differs from CleanRL, declared

* Observation statistics are updated after each iteration's learning, from
that iteration's observations, not at every step; the first iteration runs
unnormalised. CleanRL's NormalizeObservation updates at every step from
the first.
* Rewards are scaled by the return's running deviation once per iteration,
over the iteration's returns, not step by step as NormalizeReward does.
* The environments are -v5.

---

## Design

| cell | task | seeds | iterations |
| --- | --- | --- | --- |
| `ppo-cheetah` | HalfCheetah-v5 | 0, 1, 2 | 488 |
| `ppo-hopper` | Hopper-v5 | 0, 1, 2 | 488 |
| `ppo-walker` | Walker2d-v5 | 0, 1, 2 | 488 |

`gaussian-policy default ppo default` with the task's dimensions and nothing
else: rollouts of 2,048 steps, minibatches of 64, ten epochs, clip 0.2,
learning rate 3e-4 annealed to zero over the 488 iterations (999,424 steps,
CleanRL's 1,000,000 // 2,048). One environment per seed, the client
replanning every step; checkpoints every 20 iterations. E30's
`run_cell.sh`, unchanged, through `run.sh`.

Machine `guangzhao`, CPU, one thread per process, the three cells at once
beside E37; code this branch (#85 and #86 merged, plus this directory),
client plugrl-env-client at 931ab56.

---

## Checks

* **V1 - the configuration took**: every iteration's logged
`models/learning_rate` equal to (1 - (i-1)/488) x 3e-4 for iteration i,
to one part in a million. A cell that fails is not read.

---

## The status rule

As the coverage figure has it, on the mean return of iterations 479-488, on
at least **2 of 3** seeds: Hopper and Walker2d at least **500**; HalfCheetah
at least **+200** over its first iteration. A cell that passes **learns**;
one that does not **did not learn in 999,424 steps**.

---

## Predictions, and what falsifies each

**P1 - every cell runs end to end** (488 iterations logged, at least 24
checkpoints, no traceback, client exit 0, on every seed).

**P2 - it learns Hopper.** **P3 - it learns Walker2d.** **P4 - it learns
HalfCheetah.**

> Grounds for all three: known item 1 - CleanRL's returns with this
> configuration are four to five times the bars on Hopper and Walker2d, and
> 1443 on HalfCheetah, where an untrained policy starts near -300 (E6).
> For Hopper also known item 4: with the fix, 868 to 1005 by iteration 100.
> Each is falsified if fewer than two seeds clear its bar.

**Reported, not predicted:** the mean return over iterations 91-100, 241-250
and 479-488 for every seed, beside CleanRL's; wall clock.

---

## Declared deviations allowed in advance

1. One restart of any run that dies for a reason outside the experiment,
recorded in `AMENDMENT.md`.

---

## Reading order

P1, V1, the status rule, P2, P3, P4, then the reported figures.
28 changes: 28 additions & 0 deletions experiments/e38-gaussian-ppo/results/pilot.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
E38 pilot, 2026-09-27, guangzhao, code ab6e32f (#85 plus E38's scripts, before
the GAE fix), client 931ab56.

ITERS=3 SEEDS=0 bash experiments/e38-gaussian-ppo/run.sh

start 17:38:10, end 17:39:04; failed seeds: 0 in every cell.

cell seed 0 rc traceback env steps client effective_fps
ppo-cheetah 0 none 6,145 528.19
ppo-hopper 0 none 6,145 563.74
ppo-walker 0 none 6,145 584.20

Each run took about 12 s for 3 iterations of 2,048 steps including learning;
488 iterations alone would take about 30 minutes. Returns were not read.

GAE diagnostic, 2026-09-27 17:55-18:03, guangzhao. gaussian-policy under ppo
on Hopper-v5, 100 iterations of 2,048 (the schedule annealed over 100),
seeds 0-2, through E30's run_cell.sh, once in a checkout before the GAE
episode-boundary fix (ab6e32f) and once after it (274929c, #85 with #86
merged). Every run rc=0, no traceback, 100 iterations logged.

mean return seed first 41-50 91-100 length 91-100
before 0 12.1 266.5 431.5 150.4
1 15.2 588.7 1512.1 486.6
2 14.2 366.5 909.9 295.1
after 0 12.1 674.3 1005.1 312.7
1 15.2 673.1 868.0 271.4
2 14.2 453.6 949.3 296.0
11 changes: 11 additions & 0 deletions experiments/e38-gaussian-ppo/results/ppo-cheetah.out
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
cell: ppo-cheetah = gaussian-policy/default x ppo/default x HalfCheetah-v5
iters: 488 x 2048 batch: 64 replan: 1 seeds: 0 1 2
extra: algo none; policy --policy.obs-dim 17 --policy.action-dim 6
out: /home/guangzhao/zuogou/plugrl/e38/experiments/e38-gaussian-ppo/results/ppo-cheetah
start 2026-09-27 18:04:58
seed 0 finished rc=0 at 18:32:53
seed 2 finished rc=0 at 18:33:15
seed 1 finished rc=0 at 18:33:16
end 2026-09-27 18:33:16
failed seeds: 0
CELL_DONE ppo-cheetah
Loading
Loading