docs: each coverage cell shows the two commands that trained it - #15
Merged
Merged
Conversation
Clicking a cell (or opening #cov-<cell>) shows its training server and env client commands under its block, with the words that differ from the cell shown before marked, and each command copyable. Every cell names the env-client environment it needs, and the text says the part that makes the matrix possible: no server that trained these cells has MuJoCo, robosuite or gymnasium, and the clients ran in three separate environments - MuJoCo 3 for the MuJoCo tasks, robosuite 1.4.1 with MuJoCo 2.3.7 for robomimic and LIBERO. Commands come from plugrl-server's figures/coverage/commands.json.
The twelve MLP cells and their legend, rendered from this page at twice the pixel density with nothing selected. The README links it to the live figure.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The coverage figure now shows how each cell was trained, not just whether it worked.
#cov-<cell id>) and the panel under its block shows the two commands that trained it - the training server and the env client - each copyable. Words that differ from the cell shown before are marked, so going fromfpo-policy · FPOtodppo-policy · DPPOon Hopper lights up exactlydppo-policy hopper dppo hopper, the buffer, batch and iteration count, and the client's replan of 4. The MLP block opens onfpo-policy · FPO · Hopper, the pi0.5 block on its FPO cell.uv sync --extra mujoco(gymnasium with MuJoCo 3),--extra robomimicor--extra libero(robosuite 1.4.1 with MuJoCo 2.3.7).docs/media/coverage-grid.jpg: a still of the twelve MLP cells, for plugrl-server's README.Commands come from plugrl-server's
figures/coverage/commands.json(PlugRL/plugrl-server PR to follow), taken from each experiment's script; seed 0 and port 8000 throughout, output flags dropped, pi0.5's cluster paths as placeholders.mkdocs build --strictpasses; checked in a browser with#cov-dppo-dppo-hopper, where the marks and both panels render as described.