Skip to content

coverage: seven cells change - fpo-policy · DPPO, dppo-policy on Hopper, Walker2d and square learn; pi0-policy · FPO holds - #17

Merged
tactino merged 3 commits into
mainfrom
figure/coverage-green
Sep 27, 2026
Merged

tactino merged 3 commits into
mainfrom
figure/coverage-green

Conversation

@tactino

@tactino tactino commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Seven cells of the coverage figure change status, from what the experiments now show:

cell was now from
fpo-policy · DPPO · HalfCheetah no learning learns: +1,294 to +1,411 E33
fpo-policy · DPPO · Hopper no learning learns: 998 to 1,037 E33 (seeds 3-5)
fpo-policy · DPPO · Walker2d no learning learns: 283 to 569, 2 of 3 E33
dppo-policy · DPPO · Hopper still rising learns: 812 to 961 E31
dppo-policy · DPPO · Walker2d still rising learns: 528 to 622 E31
dppo-policy · DPPO · square runs learns: success 0.33-0.36 → 0.65-0.70 E34
pi0-policy · FPO collapses holds: 33 of 50 after one update E32

fpo-policy under DPPO learns once it carries dppo-policy's structure (3 x 512 layers, chunks of 4, 20 flow steps); E30 found that on Hopper, E33 confirmed it on all three tasks with the configuration fixed in advance. The cell's command panel shows those flags. dppo-policy's two MuJoCo cells needed 300 iterations instead of 100. dppo-policy on square learns when it starts, as DPPO's authors start it, from their released pretrained policy with their fine-tuning settings.

pi0-policy · FPO no longer collapses: E32 found the collapse came from how this FPO scored an action chunk, and with FPO++'s chunk loss one update leaves pi0.5 at 33 of 50. New clip of that checkpoint on the same scene as the other two; the home pages now say what was found instead of "still tracking it down", and that surviving is not yet improving.

New clips (each cell's median seed, final checkpoint, median of five evaluation episodes), coverage.json regenerated by figures/coverage/data.py, and the README still. The command panel's note no longer says every experiment ran seeds 0-2.

Merge after plugrl-server #75 (E31), #78 (E30), #79 (E33), #81 (E32) and #84 (E34): the page links each cell to its experiment's directory on main, and those five are not there yet. The matching cells.json / commands.json change is plugrl-server's figure PR.

Made through the GitHub API, not a local checkout; the files are those figures/coverage/run.sh, data.py and grid_still.py wrote on guangzhao.

…er2d learn

Five cells change from "no learning" or "still rising" to "learns":
fpo-policy under DPPO with dppo-policy's structure on HalfCheetah, Hopper and
Walker2d (plugrl-server E30, E33), and dppo-policy under DPPO on Hopper and
Walker2d at 300 iterations (E31). New clips from each cell's median seed,
coverage.json regenerated, and the grid still for the READMEs.

The command panel's note no longer says every experiment ran seeds 0-2: E33
ran Hopper on seeds 3-5.
plugrl-server E32 found why one FPO iteration took pi0.5 to zero: this FPO
scored an action chunk by averaging the error over all 320 of its elements,
most of them padding or steps never executed. Scored as FPO++ scores it, one
update leaves pi0.5 at 33 of 50. The cell moves from "collapses" to "holds",
with a new clip of that checkpoint on the same scene as the other two, and
the home pages say what was found in place of "still tracking it down" -
and that surviving is not yet improving.
@tactino tactino changed the title coverage: the fpo-policy · DPPO row and dppo-policy's Hopper and Walker2d learn coverage: fpo-policy · DPPO, dppo-policy's Hopper and Walker2d learn; pi0-policy · FPO holds Sep 27, 2026
plugrl-server E34 ran dppo-policy on square the way DPPO's authors fine-tune
it - their released checkpoint's ema weights, the last 10 of 20 denoising
steps, every value of their square config - and training success rose from
0.33-0.36 to 0.65-0.70 on every seed. New clip from the median seed's final
checkpoint, coverage.json and the README still regenerated.
@tactino tactino changed the title coverage: fpo-policy · DPPO, dppo-policy's Hopper and Walker2d learn; pi0-policy · FPO holds coverage: seven cells change - fpo-policy · DPPO, dppo-policy on Hopper, Walker2d and square learn; pi0-policy · FPO holds Sep 27, 2026
@tactino
tactino merged commit ed0d1dc into main Sep 27, 2026
1 check passed
@tactino
tactino deleted the figure/coverage-green branch September 27, 2026 18:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant