From d816d94ef2f3f539e7566bb602efff2443fcd248 Mon Sep 17 00:00:00 2001 From: tactino <18781106300@163.com> Date: Sun, 27 Sep 2026 03:43:03 -0400 Subject: [PATCH 1/2] docs: each coverage cell shows the two commands that trained it Clicking a cell (or opening #cov-) shows its training server and env client commands under its block, with the words that differ from the cell shown before marked, and each command copyable. Every cell names the env-client environment it needs, and the text says the part that makes the matrix possible: no server that trained these cells has MuJoCo, robosuite or gymnasium, and the clients ran in three separate environments - MuJoCo 3 for the MuJoCo tasks, robosuite 1.4.1 with MuJoCo 2.3.7 for robomimic and LIBERO. Commands come from plugrl-server's figures/coverage/commands.json. --- docs/index.md | 13 ++- docs/index.zh.md | 8 +- docs/javascripts/coverage.js | 128 ++++++++++++++++++++++++++++-- docs/media/coverage/coverage.json | 2 +- docs/stylesheets/coverage.css | 64 +++++++++++++++ 5 files changed, 206 insertions(+), 9 deletions(-) diff --git a/docs/index.md b/docs/index.md index 5eda63e..b44ede4 100644 --- a/docs/index.md +++ b/docs/index.md @@ -87,8 +87,10 @@ question worth answering is which combinations actually work. Below is every combination of the two MLP policies and the two algorithms on four tasks, and pi0.5 on LIBERO. The border and its label say what the experiments found. The line under each clip is the training return of all three seeds, drawn on one -scale per column, so a flat line really is flat. Hover over a cell, or tap -it, to play it. +scale per column, so a flat line really is flat. Hover over a cell to play +it; click or tap it to play it and see the two commands that trained it. +From one cell to the next, only the words that name the policy, the +algorithm and the task change.
@@ -103,6 +105,13 @@ defect of ours we are still tracking down, not a finding about FPO. The scripts that made all of this are in [figures/coverage](https://github.com/PlugRL/plugrl-server/tree/main/figures/coverage). +The two sides do not even share a Python environment. None of the servers +that trained these cells has MuJoCo, robosuite or gymnasium installed. The +env clients ran in three separate environments: gymnasium with MuJoCo 3 for +the MuJoCo tasks, and robosuite 1.4.1 with MuJoCo 2.3.7 for robomimic and for +LIBERO, because robosuite 1.4.1 does not run on MuJoCo 3. One server +codebase trained all of them. + ## A real VLA, end to end