coverage: seven cells change - fpo-policy · DPPO, dppo-policy on Hopper, Walker2d and square learn; pi0-policy · FPO holds - #17
Merged
Conversation
…er2d learn Five cells change from "no learning" or "still rising" to "learns": fpo-policy under DPPO with dppo-policy's structure on HalfCheetah, Hopper and Walker2d (plugrl-server E30, E33), and dppo-policy under DPPO on Hopper and Walker2d at 300 iterations (E31). New clips from each cell's median seed, coverage.json regenerated, and the grid still for the READMEs. The command panel's note no longer says every experiment ran seeds 0-2: E33 ran Hopper on seeds 3-5.
plugrl-server E32 found why one FPO iteration took pi0.5 to zero: this FPO scored an action chunk by averaging the error over all 320 of its elements, most of them padding or steps never executed. Scored as FPO++ scores it, one update leaves pi0.5 at 33 of 50. The cell moves from "collapses" to "holds", with a new clip of that checkpoint on the same scene as the other two, and the home pages say what was found in place of "still tracking it down" - and that surviving is not yet improving.
plugrl-server E34 ran dppo-policy on square the way DPPO's authors fine-tune it - their released checkpoint's ema weights, the last 10 of 20 denoising steps, every value of their square config - and training success rose from 0.33-0.36 to 0.65-0.70 on every seed. New clip from the median seed's final checkpoint, coverage.json and the README still regenerated.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Seven cells of the coverage figure change status, from what the experiments now show:
fpo-policy· DPPO · HalfCheetahfpo-policy· DPPO · Hopperfpo-policy· DPPO · Walker2ddppo-policy· DPPO · Hopperdppo-policy· DPPO · Walker2ddppo-policy· DPPO · squarepi0-policy· FPOfpo-policyunder DPPO learns once it carriesdppo-policy's structure (3 x 512 layers, chunks of 4, 20 flow steps); E30 found that on Hopper, E33 confirmed it on all three tasks with the configuration fixed in advance. The cell's command panel shows those flags.dppo-policy's two MuJoCo cells needed 300 iterations instead of 100.dppo-policyon square learns when it starts, as DPPO's authors start it, from their released pretrained policy with their fine-tuning settings.pi0-policy· FPO no longer collapses: E32 found the collapse came from how this FPO scored an action chunk, and with FPO++'s chunk loss one update leaves pi0.5 at 33 of 50. New clip of that checkpoint on the same scene as the other two; the home pages now say what was found instead of "still tracking it down", and that surviving is not yet improving.New clips (each cell's median seed, final checkpoint, median of five evaluation episodes),
coverage.jsonregenerated byfigures/coverage/data.py, and the README still. The command panel's note no longer says every experiment ran seeds 0-2.Merge after plugrl-server #75 (E31), #78 (E30), #79 (E33), #81 (E32) and #84 (E34): the page links each cell to its experiment's directory on
main, and those five are not there yet. The matchingcells.json/commands.jsonchange is plugrl-server's figure PR.Made through the GitHub API, not a local checkout; the files are those
figures/coverage/run.sh,data.pyandgrid_still.pywrote onguangzhao.