Skip to content

docs: each coverage cell shows the two commands that trained it - #15

Merged
tactino merged 2 commits into
mainfrom
feat/coverage-commands
Sep 27, 2026
Merged

tactino merged 2 commits into
mainfrom
feat/coverage-commands

Conversation

@tactino

@tactino tactino commented Sep 27, 2026

Copy link
Copy Markdown
Member

The coverage figure now shows how each cell was trained, not just whether it worked.

  • Click a cell (or open #cov-<cell id>) and the panel under its block shows the two commands that trained it - the training server and the env client - each copyable. Words that differ from the cell shown before are marked, so going from fpo-policy · FPO to dppo-policy · DPPO on Hopper lights up exactly dppo-policy hopper dppo hopper, the buffer, batch and iteration count, and the client's replan of 4. The MLP block opens on fpo-policy · FPO · Hopper, the pi0.5 block on its FPO cell.
  • Each command names its env-client environment: uv sync --extra mujoco (gymnasium with MuJoCo 3), --extra robomimic or --extra libero (robosuite 1.4.1 with MuJoCo 2.3.7).
  • New paragraph under the figure, both languages: none of the servers that trained these cells has MuJoCo, robosuite or gymnasium installed, and the clients ran in three separate environments, because robosuite 1.4.1 does not run on MuJoCo 3. One server codebase trained them all. Checked on each machine's venvs before writing it.
  • docs/media/coverage-grid.jpg: a still of the twelve MLP cells, for plugrl-server's README.

Commands come from plugrl-server's figures/coverage/commands.json (PlugRL/plugrl-server PR to follow), taken from each experiment's script; seed 0 and port 8000 throughout, output flags dropped, pi0.5's cluster paths as placeholders. mkdocs build --strict passes; checked in a browser with #cov-dppo-dppo-hopper, where the marks and both panels render as described.

Clicking a cell (or opening #cov-<cell>) shows its training server and env
client commands under its block, with the words that differ from the cell
shown before marked, and each command copyable. Every cell names the
env-client environment it needs, and the text says the part that makes the
matrix possible: no server that trained these cells has MuJoCo, robosuite
or gymnasium, and the clients ran in three separate environments - MuJoCo 3
for the MuJoCo tasks, robosuite 1.4.1 with MuJoCo 2.3.7 for robomimic and
LIBERO. Commands come from plugrl-server's figures/coverage/commands.json.
The twelve MLP cells and their legend, rendered from this page at twice the
pixel density with nothing selected. The README links it to the live figure.
@tactino
tactino merged commit 5a1082c into main Sep 27, 2026
1 check passed
@tactino
tactino deleted the feat/coverage-commands branch September 27, 2026 07:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant