Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
66 changes: 66 additions & 0 deletions experiments/e40-square-gaussian-ppo/FINDINGS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,66 @@
# E40: DPPO's Gaussian MLP learns robomimic square under DPPO's own Gaussian PPO, through PlugRL

2026-09-27/28 · Linux workstation (`guangzhao`), CPU only · one cell, three
seeds, 40 iterations of 80,000 steps · protocol: [`PROTOCOL.md`](PROTOCOL.md)
(`456da16`, after the pilot in `results/pilot.txt`, before the registered
run)

---

## The result

`dppo-gaussian-policy`, started from DPPO's released square Gaussian
(`square_pre_gaussian_mlp_ta4`, its `model` weights), under `ppo
dppo-square`, which is every value of DPPO's `ft_ppo_gaussian_mlp.yaml` (#92).
Training-rollout success:

| seed | start: iterations 1-2 | end: iterations 31-40 | gain |
| --- | --- | --- | --- |
| 0 | 0.185 | 0.547 | **+0.362** |
| 1 | 0.195 | 0.537 | **+0.342** |
| 2 | 0.212 | 0.565 | **+0.352** |

* **P1 holds**: 40 iterations on every seed, four checkpoints each, no
traceback, clients exited 0; 4 hours, beside E39.
* **V1 passes**: every server loaded the released checkpoint's `model`
weights, logged the deviation bounded at 1, and ran
`PPOAlgoConfigDPPOSquare`.
* **P2 holds**: it learns square, 3 of 3 seeds past +0.2.

The Gaussian · PPO row's square cell, empty until now, **learns**. That makes
Gaussian · PPO the third pair to learn all four tasks. The other two are both
DPPO pairs: `dppo-policy` · DPPO and, since E37, `fpo-policy` · DPPO.

---

## The curves

Success by window of ten iterations:

| seed | 1-10 | 11-20 | 21-30 | 31-40 |
| --- | --- | --- | --- | --- |
| 0 | 0.273 | 0.422 | 0.454 | 0.547 |
| 1 | 0.256 | 0.394 | 0.504 | 0.537 |
| 2 | 0.310 | 0.402 | 0.506 | 0.565 |

Still rising at 40. Mean return, one per step spent succeeding, went from
30-36 over the first two iterations to 109-114 over the last ten.
`approx_kl` averaged 2 to 6 x 10⁻³ and the clip fraction about 0.45
throughout. A clip of 0.01 on the ratio of the chunk's mean log-probability
binds on almost half the samples at every update, as the pilot's first
update suggested. DPPO's own setting runs that way, and it learns anyway.

---

## What E40 does not show

* **That it reproduces DPPO's curve.** These are training rollouts with the
policy's deviation, from a start of about 0.2, where DPPO plots
deterministic evaluations from about 0.3. The rates are the current
config's (1e-4), not the ones the paper's plot probably used, and GAE
here ends an episode at its time-out where DPPO's bootstraps across it.
* **Anything about E38's `gaussian-policy`.** This is DPPO's Gaussian, a
residual MLP of 1024 with chunks of 4, not CleanRL's 64 x 64.

Checkpoints and tensorboards stay on `guangzhao`. The logs, `verdicts.txt`
and `summary.tsv` are here.
133 changes: 133 additions & 0 deletions experiments/e40-square-gaussian-ppo/PROTOCOL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,133 @@
# E40 measurement protocol (pre-registered)

**Written 2026-09-27, after the pilot in `results/pilot.txt` and before the
registered run.**

This file must not be edited after the first registered data point. Anything
learned afterwards goes in `AMENDMENT.md`, dated.

---

## The question

The coverage figure's Gaussian · PPO row (E38) has no square cell. From
random weights a Gaussian has nothing to learn from on square's sparse
reward (E27). DPPO fine-tunes its own Gaussian MLP with PPO on square, from
a pretrained checkpoint it releases, as the baseline to its diffusion
policy. #92 builds that policy and that setting.

**Run the way DPPO runs it, does its Gaussian MLP learn square under PPO
through PlugRL?**

### What this cannot settle

* Whether the CleanRL-style `gaussian-policy` of E38 would learn square. This
is DPPO's Gaussian, a residual MLP of 1024 with chunks of 4, not that
one.
* Comparison with DPPO's plotted numbers, which are deterministic
evaluations under older settings (below).

---

## Declared in advance: what was already known

1. **DPPO's own Gaussian-MLP on square (state)**, from arXiv 2409.00588 Fig.
7, read off the plot and approximate: about 0.3 at the start, about 0.8
by about 5M steps, about 0.92 near 20M. These are deterministic
evaluations. They were probably produced before v0.7 of the repository,
whose config had an actor rate of 1e-5 and 1000 iterations. The config E40
runs (cc7234ad) has 1e-4 and 201 iterations.
2. **E34**: DPPO's diffusion policy in DPPO's setting, the same client
setting, 40 iterations. Training success went from 0.33-0.36 to
0.65-0.70.
3. **The released checkpoint's deviation bound is 1.0, not the config's
0.2.** The checkpoint's `logvar_max` replaces the config's value on load.
DPPO's code does this, and #92 does the same.
4. **The pilot** (`results/pilot.txt`), seed 0:
- Training success 0.250, then 0.235.
- Explained variance −0.20 after the critic-only iteration, 0.44 after
the next.
- At the first actor update, approx_kl 0.0375 and clip fraction 0.78.
- About 5.7 minutes an iteration.

---

## Design

One cell, `gauss-square`, on `guangzhao`, beside E39. Three seeds, each its
own server and two client processes of `robomimic-v1` on `square-img` (84 x
84 agentview, read by nothing). Episodes run 400 steps and do not end on
success. The client replans every 4 steps. **40 iterations** of 20,000
chunks (80,000 steps), 3.2M steps in all, as E34.

Server: `dppo-gaussian-policy default ppo dppo-square`, with
`--policy.checkpoint-path` pointing at the released checkpoint
(sha256 bf787a879dcaffe3) and nothing else changed:
- Policy: DPPO's Gaussian MLP, its `model` weights, σ from 0.1.
- Learning rates: actor 1e-4, critic 1e-3, constant.
- Minibatches of 10,000, 10 epochs.
- γ 0.999, λ 0.95, clip 0.01, target KL 1, value coefficient 0.5, no
gradient clipping.
- Rewards scaled by a 0.99 running return and clipped at 10.
- One critic-only iteration.

Not DPPO's:
- A time-out ends the episode for GAE. DPPO's GAE bootstraps across the
resets.
- There are no evaluation-only iterations.
- The data comes from 2 client processes, not 50 environments: the same
20,000 chunks per iteration.
- The environments are not reset at every iteration.

Code: this branch (#92 plus this directory). Client: plugrl-env-client at
931ab56.

---

## Checks

* **V1 - the setting took**: every seed's server log shows the released
checkpoint's `model` weights loaded with the deviation bounded at 1, and
`PPOAlgoConfigDPPOSquare`. A seed that fails is not read.

---

## The status rule

E34's, with this cell's warmup. A seed's start is its mean
`rollout/success` over iterations 1-2, which the released policy collected
unchanged. Its end is the mean over iterations 31-40. The cell **learns** if
the end exceeds the start by at least **+0.2** on at least **2 of 3**
seeds.

---

## Predictions, and what falsifies each

**P1 - the cell runs end to end**: 40 iterations logged, at least 4
checkpoints, no traceback, clients exit 0, on every seed.

**P2 - DPPO's Gaussian MLP learns square under DPPO's own Gaussian PPO.**

> Grounds: known item 1. DPPO's own curve rises from about 0.3 to about 0.8
> by 5M steps, and E40 runs 3.2M at a rate ten times the one that curve
> probably used. Known item 2: E34's policy gained 0.29-0.37 on this budget.
> Against: the pilot's first update was clipped on 78% of samples.
> Falsified if fewer than two seeds gain 0.2.

**Reported, not predicted:**
- success, approx_kl and clip fraction by window of ten iterations
- wall clock

---

## Declared deviations allowed in advance

1. One restart of any run that dies for a reason outside the experiment,
recorded in `AMENDMENT.md`.

---

## Reading order

P1, V1, the status rule, P2, then the reported figures.
13 changes: 13 additions & 0 deletions experiments/e40-square-gaussian-ppo/results/gauss-square.out
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
cell: gauss-square = dppo-gaussian-policy x ppo/dppo-square x robomimic square
iters: 40 clients per seed: 2 seeds: 0 1 2
ckpt: /home/guangzhao/zuogou/plugrl/e40-reg/../ckpt/dppo-square-gaussian/state_5000.pt sha256 bf787a879dcaffe3
server: /home/guangzhao/zuogou/plugrl/e40-reg/../plugrl-server/.venv/bin/python (src /home/guangzhao/zuogou/plugrl/e40-reg/src)
client: /home/guangzhao/zuogou/plugrl/e40-reg/../plugrl-env-client/.venv-robomimic/bin/python
out: /home/guangzhao/zuogou/plugrl/e40-reg/experiments/e40-square-gaussian-ppo/results/gauss-square
start 2026-09-27 22:28:50
seed 2 finished rc=0 at 02:30:32
seed 0 finished rc=0 at 02:30:58
seed 1 finished rc=0 at 02:31:44
end 2026-09-28 02:31:44
failed seeds: 0
CELL_DONE gauss-square
Loading
Loading