From ca6b567cc966122fdc710af130ff7d2984926488 Mon Sep 17 00:00:00 2001 From: tactino <18781106300@163.com> Date: Tue, 29 Sep 2026 15:32:33 -0400 Subject: [PATCH] docs: the guide, environment, protocol and contributing pages match the code - commands that run as written, and the env and protocol contracts in full --- docs/contributing/env.md | 41 ++++-- docs/contributing/env.zh.md | 33 +++-- docs/contributing/index.md | 16 ++- docs/contributing/index.zh.md | 14 ++- docs/contributing/pre_commit.md | 14 ++- docs/contributing/pre_commit.zh.md | 11 +- docs/env/custom_env.md | 196 ++++++++++++++++++++++++----- docs/env/custom_env.zh.md | 182 ++++++++++++++++++++++----- docs/env/index.md | 75 ++++++++--- docs/env/index.zh.md | 74 +++++++---- docs/env/libero_env.md | 22 +++- docs/env/libero_env.zh.md | 22 ++-- docs/protocol/index.md | 77 ++++++++++-- docs/protocol/index.zh.md | 58 +++++++-- docs/user_guide/get_started.md | 78 +++++++++--- docs/user_guide/get_started.zh.md | 75 +++++++---- docs/user_guide/index.md | 44 ++++--- docs/user_guide/index.zh.md | 40 +++--- 18 files changed, 828 insertions(+), 244 deletions(-) diff --git a/docs/contributing/env.md b/docs/contributing/env.md index 4c128cd..d9b9875 100644 --- a/docs/contributing/env.md +++ b/docs/contributing/env.md @@ -4,32 +4,53 @@ Environments live in `plugrl-env-client`. ## Where to edit -- Implementation: `plugrl_env_client/envs/...` -- Registration: `plugrl_env_client.utils.registration` +- Implementation: `src/plugrl_env_client/envs//_env.py`. The + CLI imports every `*_env.py` file under `envs/` at startup. +- Registration: `plugrl_env_client.utils.registration` (`register_env`, + `register_env_config`). +- Optional dependencies: a new extra under + `[project.optional-dependencies]` in `pyproject.toml`, and a check at the top + of the module that raises `ImportError` naming that extra, as + `envs/mujoco/mujoco_env.py` does. The CLI then skips the env with a warning + when the extra is missing, instead of failing to start. ## Checklist -- Implement a `BaseEnv` and a dataclass config. -- Register an env UID so `plugrl-run-env-client ` works. -- Ensure the env returns observations compatible with the worker recorder. +- Implement a `BaseEnv` and a dataclass config with a default for every field. +- Register an env UID so `plugrl-run-env-client ` works; the + subcommand is the UID in lowercase. +- Follow the [contract](../env/custom_env.md#contract): set + `single_action_space`, return batched arrays from `step`, honour + `reset_indices`, seed through `seed_rngs`, and never reset inside `step`. +- Return an `Observation` whose image and state arrays all have `num_envs` as + their leading axis. The `Recorder` (`plugrl_env_client.recorder`) slices it + per env to save first and last observations and videos. +- If the task has a notion of success, pass + `best_reward_threshold_for_success` to `register_env`, or the server's + `rollout/success` stays 0. ## Verify -Start a dummy server. +Start a dummy server, in `plugrl-server`. Set `--policy.action-dim` to your +env's action size; for a discrete action, add `--policy.discrete` and set it +to the number of choices. ```bash -plugrl-run-server dummy-policy default dummy default +uv run plugrl-run-server dummy-policy default dummy default --policy.action-dim ``` -Start an env client with your env. +Start an env client with your env, in `plugrl-env-client`. ```bash -plugrl-run-env-client --num-episodes 1 --server-host 127.0.0.1 --server-port 8000 +uv run plugrl-run-env-client --server-host 127.0.0.1 --server-port 8000 --num-episodes 3 ``` ## Troubleshooting -- Env client CLI cannot find env UID: registration module was not imported. +- Env client CLI cannot find the env UID: its module was not imported. Look + for a `Skip loading env module ...` warning at startup. +- `Expected action shape tail ...`: the dummy server's `--policy.action-dim` + does not match the env. - Env creation fails: check optional dependencies and your config defaults. ## Next steps diff --git a/docs/contributing/env.zh.md b/docs/contributing/env.zh.md index f52ddcb..d52e99c 100644 --- a/docs/contributing/env.zh.md +++ b/docs/contributing/env.zh.md @@ -4,32 +4,45 @@ ## 代码放哪里 -- 实现:`plugrl_env_client/envs/...` -- 注册:`plugrl_env_client.utils.registration` +- 实现:`src/plugrl_env_client/envs//_env.py`。CLI 启动时会 import + `envs/` 下所有 `*_env.py` 文件。 +- 注册:`plugrl_env_client.utils.registration`(`register_env`、`register_env_config`)。 +- 可选依赖:在 `pyproject.toml` 的 `[project.optional-dependencies]` 下加一个 + extra,并在模块开头检查依赖,缺了就抛出写明这个 extra 的 `ImportError`, + 和 `envs/mujoco/mujoco_env.py` 一样。这样没装 extra 时,CLI 只是打一条警告 + 跳过这个环境,而不是整个起不来。 ## 清单 -- 实现 `BaseEnv` 与配置 dataclass。 -- 注册 env UID,让 `plugrl-run-env-client ` 可用。 -- 观测格式与 recorder 兼容。 +- 实现 `BaseEnv` 与配置 dataclass,配置的每个字段都要有默认值。 +- 注册 env UID,让 `plugrl-run-env-client ` 可用;子命令是 UID 的小写形式。 +- 遵守[约定](../env/custom_env.zh.md#contract):设好 `single_action_space`, + `step` 返回成批的数组,处理 `reset_indices`,通过 `seed_rngs` 播种, + 绝不在 `step` 里自己 reset。 +- 返回的 `Observation` 里,每个图像和状态数组的第一维都是 `num_envs`。 + `Recorder`(`plugrl_env_client.recorder`)会按 env 切开它,保存首末帧观测和视频。 +- 任务有"成功"这个概念的话,给 `register_env` 传 `best_reward_threshold_for_success`, + 否则 server 的 `rollout/success` 一直是 0。 ## 验证 -启动 dummy server。 +在 `plugrl-server` 里启动 dummy server。把 `--policy.action-dim` 设成你的环境的 +动作维数;离散动作的话再加 `--policy.discrete`,并把它设成可选动作的个数。 ```bash -plugrl-run-server dummy-policy default dummy default +uv run plugrl-run-server dummy-policy default dummy default --policy.action-dim ``` -启动 env client,使用你的 env。 +在 `plugrl-env-client` 里用你的 env 启动 env client。 ```bash -plugrl-run-env-client --num-episodes 1 --server-host 127.0.0.1 --server-port 8000 +uv run plugrl-run-env-client --server-host 127.0.0.1 --server-port 8000 --num-episodes 3 ``` ## 常见问题 -- CLI 找不到 UID:注册模块没有被 import。 +- CLI 找不到 UID:模块没有被 import。看看启动时有没有 `Skip loading env module ...` 警告。 +- `Expected action shape tail ...`:dummy server 的 `--policy.action-dim` 和环境不一致。 - 创建 env 失败:检查可选依赖与默认配置。 ## 下一步 diff --git a/docs/contributing/index.md b/docs/contributing/index.md index 369c752..17b9d12 100644 --- a/docs/contributing/index.md +++ b/docs/contributing/index.md @@ -11,7 +11,9 @@ cd plugrl-server uv sync ``` -If you are working on the env client, run the same commands in `plugrl-env-client`. +If you are working on the env client, run the same commands in +`plugrl-env-client`, adding the `--extra` flags for the environments you use +(for example `uv sync --extra mujoco`). Enable pre-commit hooks. @@ -21,16 +23,20 @@ uv run pre-commit install ## Verify -Run a minimal end-to-end smoke test. +Run a minimal end-to-end smoke test. Each command runs from inside its own +repository. ```bash -# Terminal 1: plugrl-server +# Terminal 1, in plugrl-server uv run plugrl-run-server dummy-policy default dummy default -# Terminal 2: plugrl-env-client -uv run plugrl-run-env-client dummy-v1 --num-episodes 1 --server-host 127.0.0.1 --server-port 8000 +# Terminal 2, in plugrl-env-client +uv run plugrl-run-env-client dummy-v1 --server-host 127.0.0.1 --server-port 8000 --num-episodes 3 ``` +The client exits 0 after three episodes. `--server-host 127.0.0.1` is needed: +the client's default, `0.0.0.0`, is not connectable on Windows. + ## What to extend - Environments: env-client-side (`plugrl-env-client`) diff --git a/docs/contributing/index.zh.md b/docs/contributing/index.zh.md index c4ac070..d4a6742 100644 --- a/docs/contributing/index.zh.md +++ b/docs/contributing/index.zh.md @@ -11,7 +11,8 @@ cd plugrl-server uv sync ``` -如果你在改 env client 仓库,把目录换成 `plugrl-env-client`。 +如果你在改 env client,就在 `plugrl-env-client` 里跑同样的命令,并加上你要用的 +环境对应的 `--extra`(例如 `uv sync --extra mujoco`)。 启用 pre-commit。 @@ -21,16 +22,19 @@ uv run pre-commit install ## 验证 -跑一个最小端到端 smoke test。 +跑一个最小端到端 smoke test。每条命令都在它所属的仓库目录里运行。 ```bash -# 终端 1:plugrl-server +# Terminal 1, in plugrl-server uv run plugrl-run-server dummy-policy default dummy default -# 终端 2:plugrl-env-client -uv run plugrl-run-env-client dummy-v1 --num-episodes 1 --server-host 127.0.0.1 --server-port 8000 +# Terminal 2, in plugrl-env-client +uv run plugrl-run-env-client dummy-v1 --server-host 127.0.0.1 --server-port 8000 --num-episodes 3 ``` +客户端跑完三个 episode 后以 0 退出。`--server-host 127.0.0.1` 不能省:客户端的 +默认值 `0.0.0.0` 在 Windows 上连不上。 + ## 扩展点 - 环境:env client 侧,仓库为 `plugrl-env-client` diff --git a/docs/contributing/pre_commit.md b/docs/contributing/pre_commit.md index a91ba75..db74f72 100644 --- a/docs/contributing/pre_commit.md +++ b/docs/contributing/pre_commit.md @@ -12,11 +12,21 @@ uv run pre-commit install uv run pre-commit run --all-files ``` +In `plugrl-env-client`, add the `--extra` flags you use to `uv sync` +(for example `uv sync --extra mujoco`). A plain `uv sync` removes the +packages of every extra you leave out. + ## Notes - The first run may auto-fix files. Stage changes and run again. -- Formatting is handled by `ruff-format`. -- `plugrl-server` requires Python `>=3.11`. +- Linting and formatting are handled by `ruff` and `ruff-format`. +- `plugrl-server` requires Python `>=3.11,<3.14`; `plugrl-env-client` + requires `>=3.10,<3.13`. +- Each `.pre-commit-config.yaml` pins the hooks' Python with + `default_language_version`: `python3.11` in `plugrl-server`, `python3.10` in + `plugrl-env-client`. pre-commit builds the hook environments with the + `.venv`'s own Python if its version matches, and otherwise needs that + version installed somewhere it can find it. ## Next steps diff --git a/docs/contributing/pre_commit.zh.md b/docs/contributing/pre_commit.zh.md index f2f7a03..7fa286b 100644 --- a/docs/contributing/pre_commit.zh.md +++ b/docs/contributing/pre_commit.zh.md @@ -12,11 +12,18 @@ uv run pre-commit install uv run pre-commit run --all-files ``` +在 `plugrl-env-client` 里,`uv sync` 要带上你在用的 `--extra`(例如 +`uv sync --extra mujoco`)。不带的话,没点名的 extra 装的包都会被删掉。 + ## 说明 - 第一次跑可能会自动改文件。`git add` 后再跑一次。 -- 格式化由 `ruff-format` 负责。 -- `plugrl-server` 需要 Python `>=3.11`。 +- 代码检查和格式化由 `ruff` 与 `ruff-format` 负责。 +- `plugrl-server` 需要 Python `>=3.11,<3.14`;`plugrl-env-client` 需要 `>=3.10,<3.13`。 +- 两个仓库的 `.pre-commit-config.yaml` 都用 `default_language_version` 钉死了 + hook 用的 Python:`plugrl-server` 是 `python3.11`,`plugrl-env-client` 是 + `python3.10`。`.venv` 自己的 Python 版本对得上时,pre-commit 就用它建 hook 的 + 环境;对不上时,就得另外装好那个版本,并且让 pre-commit 找得到。 ## 下一步 diff --git a/docs/env/custom_env.md b/docs/env/custom_env.md index da4eea9..a0ad9d3 100644 --- a/docs/env/custom_env.md +++ b/docs/env/custom_env.md @@ -4,29 +4,46 @@ Add an env-client-side environment so `plugrl-run-env-client ` can disco ## Quickstart -Create an env class and a config dataclass, then register both. +Create an env class and a config dataclass, then register both. This one is +complete and runs as written: a point on a plane moves toward a goal, with a +2-dimensional continuous action. ```py import dataclasses + +import gymnasium as gym import numpy as np -from plugrl_env_client.envs.base_env import Action, BaseEnv, BaseEnvConfig, Observation +from plugrl_env_client.envs.base_env import ( + Action, + BaseEnv, + BaseEnvConfig, + BoolArray, + Observation, + RewardArray, +) from plugrl_env_client.utils.registration import register_env, register_env_config -UID = "custom-v1" +UID = "Point-v1" # the CLI subcommand is the lowercase form, point-v1 @register_env_config(UID) @dataclasses.dataclass -class CustomConfig(BaseEnvConfig): - ... +class PointConfig(BaseEnvConfig): + """Every field needs a default: register_env_config calls PointConfig() + at import time. Each field becomes a flag, e.g. --env.step-size.""" + + step_size: float = 0.1 # distance moved per step at full action + goal_radius: float = 0.1 # how close counts as reaching the goal + +@register_env(UID, max_episode_steps=100, best_reward_threshold_for_success=1.0) +class PointEnv(BaseEnv): + """A point on a plane moves toward a goal. Reward 1.0 on reaching it.""" -@register_env(UID) -class CustomEnv(BaseEnv): def __init__( self, - config: CustomConfig, + config: PointConfig, num_envs: int = 1, process_id: int | None = None, total_processes: int | None = None, @@ -37,57 +54,172 @@ class CustomEnv(BaseEnv): process_id=process_id, total_processes=total_processes, ) + self.step_size = config.step_size + self.goal_radius = config.goal_radius + # The env client sizes its action buffer from single_action_space + # (the action of ONE env) before the first step. + self.single_action_space = gym.spaces.Box(-1.0, 1.0, (2,), np.float32) + self.action_space = gym.spaces.Box(-1.0, 1.0, (num_envs, 2), np.float32) + self.pos = np.zeros((num_envs, 2), dtype=np.float32) + self.goal = np.zeros((num_envs, 2), dtype=np.float32) + + def _obs(self) -> Observation: + # Every array has the env batch as its leading axis. + return Observation( + images={}, + states={"pos": self.pos.copy(), "goal": self.goal.copy()}, + text="move to the goal", + ) + + def reset( + self, *, seed: int | None = None, options: dict | None = None + ) -> tuple[Observation, dict]: + self.seed_rngs(seed) # first, so that --runner.seed has an effect + idx = np.arange(self.num_envs) + if options is not None and options.get("reset_indices") is not None: + # After an episode ends the client resets only the finished envs. + idx = np.asarray(options["reset_indices"], dtype=np.int64) + # Draw from self.np_random, not np.random, or the seed does nothing. + self.pos[idx] = self.np_random.uniform(-1.0, 1.0, (len(idx), 2)) + self.goal[idx] = self.np_random.uniform(-1.0, 1.0, (len(idx), 2)) + return self._obs(), {} # the whole batch, not only the reset envs + + def step( + self, actions: Action + ) -> tuple[Observation, RewardArray, BoolArray, BoolArray, dict]: + a = np.clip(np.asarray(actions, dtype=np.float32), -1.0, 1.0) # (num_envs, 2) + self.pos = np.clip(self.pos + self.step_size * a, -1.0, 1.0) + dist = np.linalg.norm(self.pos - self.goal, axis=1) + reached = dist < self.goal_radius + reward = np.where(reached, 1.0, -dist).astype(np.float32) # (num_envs,) + terminated = reached # bool, (num_envs,) + truncated = np.zeros(self.num_envs, dtype=np.bool_) # the time limit is added for you + # Do not reset here. The client sends this terminal observation and + # then calls reset(options={"reset_indices": ...}) itself. + return self._obs(), reward, terminated, truncated, {} + + +if __name__ == "__main__": + from plugrl_env_client.cli import main + + main() +``` - def prepare_obs(self, obs: np.ndarray) -> Observation: - return Observation(images={}, states={}, text="") +## Where the file goes - def reset(self, *, seed: int | None = None, options: dict | None = None) -> tuple[Observation | None, dict]: - ... +The CLI can only offer an env whose module was imported before the CLI was +built. There are two ways to get there. - def step(self, action: Action) -> tuple[Observation | None, float, bool, bool, dict]: - ... +**Inside the env client.** Save the file as +`src/plugrl_env_client/envs/point/point_env.py` in your `plugrl-env-client` +checkout, with an empty `__init__.py` beside it like the shipped families. +At startup the CLI imports every file named `*_env.py` under +`plugrl_env_client/envs/`, so `uv run plugrl-run-env-client point-v1` finds +it. If that import fails, the CLI only prints a `Skip loading env module ...` +warning and the ID is missing. + +**Anywhere else.** Keep the `if __name__ == "__main__":` block and run the +file itself with the env client's Python, for example from inside +`plugrl-env-client`: + +```bash +uv run python /path/to/point_env.py point-v1 --help ``` +Importing the file registers the env, and `plugrl_env_client.cli.main()` then +builds the CLI with it included. `plugrl-run-env-client` on its own never +imports your file, so it will not list the env. The CLI module builds its +list of envs when it is first imported, so if you write your own launcher, +import your env module before `plugrl_env_client.cli`. +`examples/pusht/pusht_env.py` in `plugrl-env-client` is built this way. + ## Verify -Env should appear as a CLI subcommand. +The env should appear as a CLI subcommand (in-tree placement shown; for the +other, replace `plugrl-run-env-client` with `python /path/to/point_env.py`): ```bash -plugrl-run-env-client custom-v1 --help +uv run plugrl-run-env-client point-v1 --help ``` -Run one episode against a dummy server. +Run it against a dummy server. The dummy policy's action is 7-dimensional by +default, for `dummy-v1`, so match this env's 2 dimensions: ```bash -plugrl-run-server dummy-policy default dummy default -plugrl-run-env-client custom-v1 --num-episodes 1 +# Terminal A, in plugrl-server +uv run plugrl-run-server dummy-policy default dummy default --policy.action-dim 2 + +# Terminal B, in plugrl-env-client +uv run plugrl-run-env-client point-v1 --server-host 127.0.0.1 --server-port 8000 --num-episodes 3 ``` +The client exits 0. `runs//rollout/proc_000/summary.json` has the +episode count and mean return. + ## Contract -- Env inherits `BaseEnv`. -- Config inherits `BaseEnvConfig`. -- Implement `reset` and `step`. -- Convert raw env outputs into `Observation` in `prepare_obs`. +- Env inherits `BaseEnv`. Config inherits `BaseEnvConfig` and is a dataclass + whose every field has a default. - `__init__` takes `config, num_envs, process_id, total_processes`, the same four as `BaseEnv.__init__` and as the shipped `MuJoCoEnv`. `EnvSpec.make` always passes `num_envs`, and `gym.make_vec` forwards `process_id` and - `total_processes`. An earlier version of this page used `worker_id` and - `total_workers`; those names appear nowhere in `plugrl-env-client`, and a - class with that signature raises `TypeError` on the unexpected `num_envs`. + `total_processes`. Both are `None` unless the client runs with + `--runner.pass-proc-id`. An earlier version of this page used `worker_id` + and `total_workers`; those names appear nowhere in `plugrl-env-client`, and + a class with that signature raises `TypeError` on the unexpected `num_envs`. +- `__init__` sets `single_action_space`, the space of one env's action. The + client reads its shape and dtype before the first step; `action_space` is + only a fallback. +- Everything is batched over `num_envs`. `step` receives actions of shape + `(num_envs, *action_shape)` and returns `(Observation, reward, terminated, + truncated, info)`: reward a float32 array of shape `(num_envs,)`, + terminated and truncated bool arrays of shape `(num_envs,)`, info a dict. + Every image and state array in the `Observation` has `num_envs` as its + leading axis; `text` may be a single string. +- `reset(*, seed=None, options=None)` returns `(Observation, info)` for the + whole batch. When episodes end, the client calls + `reset(options={"reset_indices": ...})`: reset only those envs, and still + return the observation of all of them. The option arrives with + `num_envs=1` too, so do not pass it on to a wrapped Gymnasium env. +- `reset` calls `self.seed_rngs(seed)` first and draws all randomness from + `self.np_random`. Otherwise `--runner.seed` does nothing. The client passes + the seed on the first reset only. +- `step` must not reset the env itself. The step that reports terminated or + truncated returns that episode's last observation; the client sends it as + the terminal observation and resets the env afterwards + ([SPEC.md §5.4](https://github.com/PlugRL/plugrl-protocol/blob/main/SPEC.md)). + An env that resets inside `step` sends the next episode's first + observation instead, and nothing raises. +- info can be `{}`. The client adds the `episode` statistics the server + reads; the server reads nothing else from info. ## Registration -- `register_env_config` registers the config dataclass. -- `register_env` registers the env class. -- `register_env` supports optional parameters such as `max_episode_steps`. +- `register_env_config(uid)` registers the config dataclass. It calls the + class with no arguments at import time, which is why every field needs a + default. +- `register_env(uid, ...)` registers the env class. The CLI subcommand is + `uid` in lowercase. +- `max_episode_steps` adds a time limit that sets `truncated`. + `--runner.max-episode-steps` overrides it at run time. +- `best_reward_threshold_for_success`: an episode counts as a success if any + of its step rewards reaches this value. That decides the `s` the server + averages into `rollout/success`. Without it, success is always false and + `rollout/success` stays 0. +- Any other keyword argument is passed to the env's `__init__`, and must be + JSON-serializable, or `register_env` raises `RuntimeError`. ## Troubleshooting -- Env ID not listed: module import did not run. -- Multi process init conflicts: try `--runner.use-env-lock`. +| Symptom | Cause | +|---|---| +| Env ID not listed | Its module was not imported. In-tree: the file name must end in `_env.py`; look for a `Skip loading env module` warning. Elsewhere: run the file itself, not `plugrl-run-env-client` | +| `Env must expose action space via single_action_space/action_space` | `__init__` did not set `single_action_space` | +| `Expected action shape tail (2,), got (7,)` | The server policy's action dimension differs from the env's; set `--policy.action-dim` on the server | +| `Expected reward shape (1,), got ()` | `step` returned a plain float; return an array of shape `(num_envs,)` | +| Multi process init conflicts | Try `--runner.use-env-lock` | ## Next steps - [Environments](index.md) -- [Get Started](../user_guide/get_started.md) \ No newline at end of file +- [Get Started](../user_guide/get_started.md) diff --git a/docs/env/custom_env.zh.md b/docs/env/custom_env.zh.md index 91905c4..a380199 100644 --- a/docs/env/custom_env.zh.md +++ b/docs/env/custom_env.zh.md @@ -4,30 +4,45 @@ ## 快速开始 -实现一个 env 类与一个 config dataclass,并注册它们。 +实现一个 env 类与一个 config dataclass,并注册它们。下面这个是完整的,照抄就能跑: +平面上一个点朝目标移动,动作是 2 维连续量。 ```py import dataclasses +import gymnasium as gym import numpy as np -from plugrl_env_client.envs.base_env import Action, BaseEnv, BaseEnvConfig, Observation +from plugrl_env_client.envs.base_env import ( + Action, + BaseEnv, + BaseEnvConfig, + BoolArray, + Observation, + RewardArray, +) from plugrl_env_client.utils.registration import register_env, register_env_config -UID = "custom-v1" +UID = "Point-v1" # the CLI subcommand is the lowercase form, point-v1 @register_env_config(UID) @dataclasses.dataclass -class CustomConfig(BaseEnvConfig): - ... +class PointConfig(BaseEnvConfig): + """Every field needs a default: register_env_config calls PointConfig() + at import time. Each field becomes a flag, e.g. --env.step-size.""" + step_size: float = 0.1 # distance moved per step at full action + goal_radius: float = 0.1 # how close counts as reaching the goal + + +@register_env(UID, max_episode_steps=100, best_reward_threshold_for_success=1.0) +class PointEnv(BaseEnv): + """A point on a plane moves toward a goal. Reward 1.0 on reaching it.""" -@register_env(UID) -class CustomEnv(BaseEnv): def __init__( self, - config: CustomConfig, + config: PointConfig, num_envs: int = 1, process_id: int | None = None, total_processes: int | None = None, @@ -38,56 +53,157 @@ class CustomEnv(BaseEnv): process_id=process_id, total_processes=total_processes, ) + self.step_size = config.step_size + self.goal_radius = config.goal_radius + # The env client sizes its action buffer from single_action_space + # (the action of ONE env) before the first step. + self.single_action_space = gym.spaces.Box(-1.0, 1.0, (2,), np.float32) + self.action_space = gym.spaces.Box(-1.0, 1.0, (num_envs, 2), np.float32) + self.pos = np.zeros((num_envs, 2), dtype=np.float32) + self.goal = np.zeros((num_envs, 2), dtype=np.float32) + + def _obs(self) -> Observation: + # Every array has the env batch as its leading axis. + return Observation( + images={}, + states={"pos": self.pos.copy(), "goal": self.goal.copy()}, + text="move to the goal", + ) + + def reset( + self, *, seed: int | None = None, options: dict | None = None + ) -> tuple[Observation, dict]: + self.seed_rngs(seed) # first, so that --runner.seed has an effect + idx = np.arange(self.num_envs) + if options is not None and options.get("reset_indices") is not None: + # After an episode ends the client resets only the finished envs. + idx = np.asarray(options["reset_indices"], dtype=np.int64) + # Draw from self.np_random, not np.random, or the seed does nothing. + self.pos[idx] = self.np_random.uniform(-1.0, 1.0, (len(idx), 2)) + self.goal[idx] = self.np_random.uniform(-1.0, 1.0, (len(idx), 2)) + return self._obs(), {} # the whole batch, not only the reset envs + + def step( + self, actions: Action + ) -> tuple[Observation, RewardArray, BoolArray, BoolArray, dict]: + a = np.clip(np.asarray(actions, dtype=np.float32), -1.0, 1.0) # (num_envs, 2) + self.pos = np.clip(self.pos + self.step_size * a, -1.0, 1.0) + dist = np.linalg.norm(self.pos - self.goal, axis=1) + reached = dist < self.goal_radius + reward = np.where(reached, 1.0, -dist).astype(np.float32) # (num_envs,) + terminated = reached # bool, (num_envs,) + truncated = np.zeros(self.num_envs, dtype=np.bool_) # the time limit is added for you + # Do not reset here. The client sends this terminal observation and + # then calls reset(options={"reset_indices": ...}) itself. + return self._obs(), reward, terminated, truncated, {} + + +if __name__ == "__main__": + from plugrl_env_client.cli import main + + main() +``` + +## 文件放在哪里 - def prepare_obs(self, obs: np.ndarray) -> Observation: - return Observation(images={}, states={}, text="") +CLI 只能提供那些在它构建之前模块就已被 import 的环境。有两种办法做到。 - def reset(self, *, seed: int | None = None, options: dict | None = None) -> tuple[Observation | None, dict]: - ... +**放进 env client 里。** 在你的 `plugrl-env-client` checkout 里把文件存成 +`src/plugrl_env_client/envs/point/point_env.py`,旁边放一个空的 `__init__.py`, +和内置的各个家族一样。CLI 启动时会 import `plugrl_env_client/envs/` 下所有名为 +`*_env.py` 的文件,所以 `uv run plugrl-run-env-client point-v1` 能找到它。 +如果这次 import 失败,CLI 只会打印一条 `Skip loading env module ...` 警告, +这个 ID 就不见了。 - def step(self, action: Action) -> tuple[Observation | None, float, bool, bool, dict]: - ... +**放在别处。** 保留文件末尾的 `if __name__ == "__main__":`,用 env client 的 +Python 直接运行这个文件,比如在 `plugrl-env-client` 目录里: + +```bash +uv run python /path/to/point_env.py point-v1 --help ``` +import 这个文件就完成了注册,随后 `plugrl_env_client.cli.main()` 构建出的 CLI +里就有它。单独运行 `plugrl-run-env-client` 永远不会 import 你的文件,所以列不出 +这个环境。CLI 模块在第一次被 import 时就定下环境列表,所以如果你自己写启动脚本, +要先 import 你的环境模块,再 import `plugrl_env_client.cli`。`plugrl-env-client` +里的 `examples/pusht/pusht_env.py` 就是这么写的。 + ## 验证 -注册后应该能看到子命令。 +注册后应该能看到子命令(这里是放进 env client 的写法;放在别处时把 +`plugrl-run-env-client` 换成 `python /path/to/point_env.py`): ```bash -plugrl-run-env-client custom-v1 --help +uv run plugrl-run-env-client point-v1 --help ``` -对着 dummy server 跑一个 episode。 +对着 dummy server 跑一下。dummy 策略的动作默认是 7 维,那是给 `dummy-v1` 的, +这里要改成和这个环境一样的 2 维: ```bash -plugrl-run-server dummy-policy default dummy default -plugrl-run-env-client custom-v1 --num-episodes 1 +# Terminal A, in plugrl-server +uv run plugrl-run-server dummy-policy default dummy default --policy.action-dim 2 + +# Terminal B, in plugrl-env-client +uv run plugrl-run-env-client point-v1 --server-host 127.0.0.1 --server-port 8000 --num-episodes 3 ``` -## 约定 +客户端以 0 退出。`runs//rollout/proc_000/summary.json` 里有 episode 数 +和平均回报。 + +## 约定 {#contract} -- env 继承 `BaseEnv` -- config 继承 `BaseEnvConfig` -- 实现 `reset` 与 `step` -- 在 `prepare_obs` 中把原始输出转成 `Observation` +- env 继承 `BaseEnv`。config 继承 `BaseEnvConfig`,是一个 dataclass,每个字段都要有默认值。 - `__init__` 接收 `config, num_envs, process_id, total_processes` 四个参数, 与 `BaseEnv.__init__` 以及内置的 `MuJoCoEnv` 一致。`EnvSpec.make` 总会传 - `num_envs`,`gym.make_vec` 会转发 `process_id` 与 `total_processes`。本页 - 此前用的是 `worker_id` 与 `total_workers`,这两个名字在 `plugrl-env-client` - 里根本不存在,按那个签名写的类会因为多出来的 `num_envs` 直接抛 `TypeError`。 + `num_envs`,`gym.make_vec` 会转发 `process_id` 与 `total_processes`;除非客户端 + 带了 `--runner.pass-proc-id`,这两个都是 `None`。本页此前用的是 `worker_id` 与 + `total_workers`,这两个名字在 `plugrl-env-client` 里根本不存在,按那个签名写的类 + 会因为多出来的 `num_envs` 直接抛 `TypeError`。 +- `__init__` 里要设 `single_action_space`,即单个 env 的动作空间。客户端在第一步 + 之前就读它的 shape 和 dtype;`action_space` 只是备选。 +- 一切都按 `num_envs` 成批。`step` 收到的动作形状是 `(num_envs, *action_shape)`, + 返回 `(Observation, reward, terminated, truncated, info)`:reward 是形状 + `(num_envs,)` 的 float32 数组,terminated 和 truncated 是形状 `(num_envs,)` 的 + bool 数组,info 是 dict。`Observation` 里每个图像和状态数组的第一维都是 + `num_envs`;`text` 可以只给一个字符串。 +- `reset(*, seed=None, options=None)` 返回整批的 `(Observation, info)`。有 episode + 结束时,客户端会调用 `reset(options={"reset_indices": ...})`:只重置这些 env, + 但返回的仍是全部 env 的观测。`num_envs=1` 时这个选项照样会传进来,所以别把它 + 转交给内层的 Gymnasium 环境。 +- `reset` 一开头先调 `self.seed_rngs(seed)`,所有随机数都从 `self.np_random` 取。 + 否则 `--runner.seed` 不起作用。客户端只在第一次 reset 时传种子。 +- `step` 里不能自己 reset。报告 terminated 或 truncated 的那一步,要返回这个 + episode 的最后一帧观测;客户端把它当作终止观测发出去,之后再重置这个 env + ([SPEC.md §5.4](https://github.com/PlugRL/plugrl-protocol/blob/main/SPEC.md))。 + 在 `step` 里自己 reset 的环境,发出去的会是下一个 episode 的第一帧,而且不会报任何错。 +- info 可以是 `{}`。server 要读的 `episode` 统计由客户端自己加上;info 里的其他 + 内容 server 一概不读。 ## 注册 -- `register_env_config` 注册 config dataclass -- `register_env` 注册 env 类 -- `register_env` 支持 `max_episode_steps` 等可选参数 +- `register_env_config(uid)` 注册 config dataclass。它在 import 时不带参数地实例化 + 这个类,所以每个字段都要有默认值。 +- `register_env(uid, ...)` 注册 env 类。CLI 子命令是 `uid` 的小写形式。 +- `max_episode_steps` 会加一个时间上限,到点时置 `truncated`。运行时可用 + `--runner.max-episode-steps` 覆盖。 +- `best_reward_threshold_for_success`:一个 episode 里只要有一步奖励达到这个值, + 就算成功。server 平均进 `rollout/success` 的 `s` 就由它决定。不设的话成功永远是 + false,`rollout/success` 一直是 0。 +- 其他关键字参数会传给 env 的 `__init__`,而且必须能 JSON 序列化,否则 + `register_env` 抛 `RuntimeError`。 ## 常见问题 -- CLI 找不到 env id:模块没有被 import。 -- 多进程初始化冲突:可尝试 `--runner.use-env-lock`。 +| 现象 | 原因 | +|---|---| +| CLI 里没有这个 env ID | 模块没被 import。放在 env client 里的:文件名要以 `_env.py` 结尾,找找有没有 `Skip loading env module` 警告。放在别处的:直接运行这个文件,而不是 `plugrl-run-env-client` | +| `Env must expose action space via single_action_space/action_space` | `__init__` 里没设 `single_action_space` | +| `Expected action shape tail (2,), got (7,)` | server 策略的动作维数和环境的不一致;在 server 上设 `--policy.action-dim` | +| `Expected reward shape (1,), got ()` | `step` 返回了一个普通 float;要返回形状 `(num_envs,)` 的数组 | +| 多进程初始化冲突 | 试试 `--runner.use-env-lock` | ## 下一步 - [环境概览](index.zh.md) -- [快速开始](../user_guide/get_started.zh.md) \ No newline at end of file +- [快速开始](../user_guide/get_started.zh.md) diff --git a/docs/env/index.md b/docs/env/index.md index 1963da3..10a8652 100644 --- a/docs/env/index.md +++ b/docs/env/index.md @@ -4,23 +4,26 @@ Environments run on the env client and are created via Gymnasium. ## Quickstart -List options for one environment. +List the options for one environment, from inside `plugrl-env-client`. ```bash -plugrl-run-env-client dummy-v1 --help +uv run plugrl-run-env-client dummy-v1 --help ``` -Run one short episode. +Run a few short episodes against a dummy server. ```bash -plugrl-run-server dummy-policy default dummy default -plugrl-run-env-client dummy-v1 --num-episodes 1 --server-host 127.0.0.1 --server-port 8000 +# Terminal A, in plugrl-server +uv run plugrl-run-server dummy-policy default dummy default + +# Terminal B, in plugrl-env-client +uv run plugrl-run-env-client dummy-v1 --server-host 127.0.0.1 --server-port 8000 --num-episodes 3 ``` ## Verify -- Env client prints server metadata. -- Env client resets and steps the environment. +- The env client prints `Server metadata: {...}`. +- It resets and steps the environment, and exits 0 after the episodes. ## How environments are created @@ -44,28 +47,66 @@ env = gym.make_vec( ` registered but entry_point is not specified`. This page previously showed the `gym.make` form; that form never worked. +The CLI finds an environment only if its module was imported before the CLI +was built. It imports every `*_env.py` file under `plugrl_env_client/envs/` +itself; code anywhere else has to import itself first. Both ways are in +[Custom environment](custom_env.md). + ## Built-in environment IDs -- `dummy-v1` -- `mujoco-v1` - needs the `mujoco` extra; this is the env the quickstart uses -- `classic-v1` -- `atari-v1` -- `robomimic-v1` -- `d4rl-*` when optional deps are installed -- `libero-*` when optional deps are installed +Only `dummy-v1` works on a plain `uv sync`. Every other family needs its +extra, e.g. `uv sync --extra classic`; without it the ID is missing from the +CLI and startup prints a `Skip loading env module ...` warning naming the +extra. + +| Env ID | Extra | What it runs | +|---|---|---| +| `dummy-v1` | none | Random images, states and rewards. For connectivity checks | +| `mujoco-v1` | `mujoco` | Gymnasium's MuJoCo tasks, `HalfCheetah-v5` by default (`--env.name`). The quickstart env | +| `classic-v1` | `classic` | Gymnasium's classic control, `CartPole-v1` by default | +| `atari-v1` | `atari` | ALE Atari games, `BreakoutNoFrameskip-v4` by default | +| `d4rl-v1` | `d4rl` | D4RL's MuJoCo tasks, `hopper-medium-v2` by default | +| `robomimic-v1` | `robomimic` | robomimic's robosuite tasks; see below | +| `libero-v1` | `libero` | LIBERO task suites; see [Libero environment](libero_env.md) | + +`robomimic-v1` takes `--env.name` from `lift`, `can`, `square` and +`transport`, each also as an `-img` variant (default `can-img`). The `-img` +variants add the wrist cameras' images to the observation; the others carry +only the rendered `agentview` frame beside the states. An episode +ends as terminated on the step the task succeeds (turn that off with +`--env.no-terminate-on-success`), and is truncated at `--env.horizon`, which +defaults to robomimic's own rollout horizon: 400 steps for lift, can and +square, 700 for transport. It runs one env per process, and it refuses +`--runner.seed`, because the robosuite simulation underneath is not seeded. +The `robomimic` extra pins MuJoCo 2.3.7 and robosuite 1.4.1, so it needs an +environment of its own, apart from the `mujoco` extra; see +[Libero environment](libero_env.md#install). It also builds `egl-probe` with +CMake, so install `cmake` first. ## Common env client flags +- `--num-envs`: environments per client process - `--num-procs`: run multiple env client processes -- `--server-host`, `--server-port`: server address -- `--recorder.video-fps`: output fps for recorded mp4 artifacts +- `--server-host`, `--server-port`: server address. Always pass + `--server-host`; its default, `0.0.0.0`, is not connectable on Windows +- `--runner.seed`: base seed; without it episodes differ between runs +- `--exp-name`: names the output directory, `runs//` + +Recording is off by default. `--recorder.episode-freq N` records every Nth +finished episode: its first and last observation go under +`runs//rollout/proc_000/sampled/`. Add `--recorder.record-video` to +also write an mp4 per image key when the recorded episode is env 0's; +`--recorder.video-fps` sets its frame rate. Videos need ffmpeg, which the +base install leaves out: add the `video` extra (`uv sync --extra video`, +alongside the others). Only process 0 records unless you pass +`--recorder.no-thread0-only`. There is no flag for running an env at a fixed wall-clock FPS. `--use-real-time` and `--fps` were listed here and do not exist on this CLI. ## Troubleshooting -- Env ID not found in CLI: registration module was not imported. +- Env ID not found in CLI: its extra is not installed, or its module was not imported. - Multi process init conflicts: try `--runner.use-env-lock` if your env is heavy. ## Next steps diff --git a/docs/env/index.zh.md b/docs/env/index.zh.md index 7f27452..2f2aef9 100644 --- a/docs/env/index.zh.md +++ b/docs/env/index.zh.md @@ -2,27 +2,28 @@ 环境运行在 env client 侧,通过 Gymnasium 创建。 -> Note: 环境代码可以放在你自己的包里。只要 env client 启动时能 import 并完成注册,CLI 就能发现它。 - ## 快速开始 -查看一个环境的参数。 +在 `plugrl-env-client` 里查看一个环境的参数。 ```bash -plugrl-run-env-client dummy-v1 --help +uv run plugrl-run-env-client dummy-v1 --help ``` -跑一个短 episode。 +对着 dummy server 跑几个短 episode。 ```bash -plugrl-run-server dummy-policy default dummy default -plugrl-run-env-client dummy-v1 --num-episodes 1 --server-host 127.0.0.1 --server-port 8000 +# Terminal A, in plugrl-server +uv run plugrl-run-server dummy-policy default dummy default + +# Terminal B, in plugrl-env-client +uv run plugrl-run-env-client dummy-v1 --server-host 127.0.0.1 --server-port 8000 --num-episodes 3 ``` ## 验证 -- env client 打印 server 元信息 -- env client 能 reset 与 step 环境 +- env client 打印 `Server metadata: {...}` +- 它能 reset、step 环境,跑完这些 episode 后以 0 退出 ## 环境如何创建 @@ -46,28 +47,59 @@ env = gym.make_vec( ` registered but entry_point is not specified`。本页此前写的是 `gym.make` 形式,那个写法从来跑不通。 -## 常见内置环境 ID - -- `dummy-v1` -- `mujoco-v1`:需要 `mujoco` 可选依赖;快速开始用的就是它 -- `classic-v1` -- `atari-v1` -- `robomimic-v1` -- `d4rl-*` 需要安装可选依赖 -- `libero-*` 需要安装可选依赖 +CLI 只认得在它构建之前就已经被 import 过的环境模块。`plugrl_env_client/envs/` +下所有 `*_env.py` 文件由它自己 import;放在别处的代码得自己先 import 自己。 +两种做法都写在[自定义环境](custom_env.zh.md)里。 + +## 内置环境 ID + +普通的 `uv sync` 之后只有 `dummy-v1` 能用。其余每个家族都要装对应的 extra, +例如 `uv sync --extra classic`;没装的话 CLI 里没有这个 ID,启动时会打印一条 +`Skip loading env module ...` 警告,里面写着缺哪个 extra。 + +| Env ID | Extra | 跑的是什么 | +|---|---|---| +| `dummy-v1` | 无 | 随机的图像、状态和奖励,用来检查连通性 | +| `mujoco-v1` | `mujoco` | Gymnasium 的 MuJoCo 任务,默认 `HalfCheetah-v5`(`--env.name`)。快速开始用的就是它 | +| `classic-v1` | `classic` | Gymnasium 的 classic control,默认 `CartPole-v1` | +| `atari-v1` | `atari` | ALE 的 Atari 游戏,默认 `BreakoutNoFrameskip-v4` | +| `d4rl-v1` | `d4rl` | D4RL 的 MuJoCo 任务,默认 `hopper-medium-v2` | +| `robomimic-v1` | `robomimic` | robomimic 的 robosuite 任务,见下文 | +| `libero-v1` | `libero` | LIBERO 任务套件,见 [Libero 环境](libero_env.zh.md) | + +`robomimic-v1` 的 `--env.name` 可选 `lift`、`can`、`square`、`transport`,每个 +还有 `-img` 版本(默认 `can-img`)。`-img` 版本会把腕部相机的图像放进观测;不带 +`-img` 的只有状态,外加单独渲染的 `agentview` 画面。任务成功的那一步,episode 以 terminated 结束 +(用 `--env.no-terminate-on-success` 关掉);到 `--env.horizon` 步时被截断,默认 +取 robomimic 自己的 rollout 长度:lift、can、square 为 400 步,transport 为 700 步。 +它每个进程只跑一个 env,并且拒绝 `--runner.seed`,因为底下的 robosuite 仿真 +没有被播种。`robomimic` extra 钉死了 MuJoCo 2.3.7 和 robosuite 1.4.1,所以要 +单独一个环境,不能和 `mujoco` extra 装在一起,见 +[Libero 环境](libero_env.zh.md#install)。它还要用 CMake 编译 `egl-probe`, +先装好 `cmake`。 ## 常用 env client 参数 +- `--num-envs`:每个客户端进程里的环境数 - `--num-procs`:多进程并行跑环境 -- `--server-host`、`--server-port`:server 地址 -- `--recorder.video-fps`:录制 mp4 的输出帧率 +- `--server-host`、`--server-port`:server 地址。`--server-host` 一定要传, + 默认的 `0.0.0.0` 在 Windows 上连不上 +- `--runner.seed`:基础种子;不设的话每次运行的 episode 都不一样 +- `--exp-name`:输出目录 `runs//` 的名字 + +默认不录制。`--recorder.episode-freq N` 每结束 N 个 episode 录一个:它的第一帧 +和最后一帧观测存到 `runs//rollout/proc_000/sampled/` 下。再加 +`--recorder.record-video`,当被录的是 0 号 env 的 episode 时,每个图像键还会写 +一个 mp4;帧率用 `--recorder.video-fps` 设。写视频要用 ffmpeg,基础安装不带它: +需要加上 `video` extra(`uv sync --extra video`,和其他 extra 一起写)。默认只有 0 号进程录制,加 +`--recorder.no-thread0-only` 让每个进程都录。 没有让环境按固定墙钟 FPS 运行的参数。本页此前列出的 `--use-real-time`、 `--fps` 在这个 CLI 上并不存在。 ## 常见问题 -- CLI 找不到 env id:注册模块没有被 import。 +- CLI 找不到 env ID:对应的 extra 没装,或者模块没有被 import。 - 多进程初始化冲突:环境较重时可尝试 `--runner.use-env-lock`。 ## 下一步 diff --git a/docs/env/libero_env.md b/docs/env/libero_env.md index 285dbc2..4959cd6 100644 --- a/docs/env/libero_env.md +++ b/docs/env/libero_env.md @@ -4,24 +4,34 @@ Run Libero tasks in `plugrl-env-client` via an optional dependency group. ## Install -Install env client with Libero extras. +LIBERO runs on robosuite 1.4.1, which needs MuJoCo 2.3.7: MuJoCo 3 fails +robosuite 1.4.1's joint-type assertion. `mujoco-v1` and the other Gymnasium +families are meant for MuJoCo 3. So give LIBERO an environment of its own, +for example a second clone of the env client. `robomimic-v1` uses the same +MuJoCo 2.3.7 / robosuite 1.4.1 stack and needs the same separation. ```bash -pip install -e ".[libero]" +git clone https://github.com/PlugRL/plugrl-env-client.git plugrl-env-client-libero +cd plugrl-env-client-libero +uv sync --extra libero +uv run python -c "import mujoco, robosuite; print(mujoco.__version__, robosuite.__version__)" ``` +The last line should print `2.3.7 1.4.1`. Run the commands below from inside +this clone. + ## Quickstart Inspect the CLI config. ```bash -plugrl-run-env-client libero-v1 --help +uv run plugrl-run-env-client libero-v1 --help ``` Run a few episodes. ```bash -plugrl-run-env-client libero-v1 --num-episodes 10 --server-host 127.0.0.1 --server-port 8000 +uv run plugrl-run-env-client libero-v1 --num-episodes 10 --server-host 127.0.0.1 --server-port 8000 ``` ## As E11 ran it @@ -31,7 +41,7 @@ client process per task, ten tasks at once, with each task's initial states taken in order rather than sampled. ```bash -plugrl-run-env-client libero-v1 \ +uv run plugrl-run-env-client libero-v1 \ --server-host 127.0.0.1 --server-port 8000 \ --num-envs 1 --num-procs 10 --num-episodes 10 \ --env.task-suite-name libero_spatial \ @@ -66,7 +76,7 @@ Protocol, results and the recorded environment of both processes: ## Troubleshooting -- Multi process init conflicts: try `--use-env-lock`. +- Multi process init conflicts: try `--runner.use-env-lock`. ## Next steps diff --git a/docs/env/libero_env.zh.md b/docs/env/libero_env.zh.md index 7185d78..9addce0 100644 --- a/docs/env/libero_env.zh.md +++ b/docs/env/libero_env.zh.md @@ -2,26 +2,34 @@ 通过可选依赖在 `plugrl-env-client` 中运行 Libero。 -## 安装 +## 安装 {#install} -安装带 Libero extras 的 env client。 +LIBERO 跑在 robosuite 1.4.1 上,而 robosuite 1.4.1 需要 MuJoCo 2.3.7:换成 +MuJoCo 3 会在 robosuite 1.4.1 的关节类型断言上失败。`mujoco-v1` 和其他 Gymnasium +家族按 MuJoCo 3 来用。所以给 LIBERO 单独一个环境,比如再 clone 一份 env client。 +`robomimic-v1` 用的是同一套 MuJoCo 2.3.7 / robosuite 1.4.1,也要同样分开。 ```bash -pip install -e ".[libero]" +git clone https://github.com/PlugRL/plugrl-env-client.git plugrl-env-client-libero +cd plugrl-env-client-libero +uv sync --extra libero +uv run python -c "import mujoco, robosuite; print(mujoco.__version__, robosuite.__version__)" ``` +最后一行应该打印 `2.3.7 1.4.1`。下面的命令都在这份 clone 里运行。 + ## 快速开始 查看可配置项。 ```bash -plugrl-run-env-client libero-v1 --help +uv run plugrl-run-env-client libero-v1 --help ``` 跑几个 episode。 ```bash -plugrl-run-env-client libero-v1 --num-episodes 10 --server-host 127.0.0.1 --server-port 8000 +uv run plugrl-run-env-client libero-v1 --num-episodes 10 --server-host 127.0.0.1 --server-port 8000 ``` ## E11 是怎么跑的 @@ -30,7 +38,7 @@ E11 通过一个 PlugRL server 在 LIBERO 上评测 `pi05_libero` checkpoint: 客户端进程,十个任务同时跑,每个任务的初始状态按顺序取,而不是随机采样。 ```bash -plugrl-run-env-client libero-v1 \ +uv run plugrl-run-env-client libero-v1 \ --server-host 127.0.0.1 --server-port 8000 \ --num-envs 1 --num-procs 10 --num-episodes 10 \ --env.task-suite-name libero_spatial \ @@ -62,7 +70,7 @@ plugrl-run-env-client libero-v1 \ ## 常见问题 -- 多进程初始化冲突:可尝试 `--use-env-lock`。 +- 多进程初始化冲突:可尝试 `--runner.use-env-lock`。 ## 下一步 diff --git a/docs/protocol/index.md b/docs/protocol/index.md index 05109f7..ce933fb 100644 --- a/docs/protocol/index.md +++ b/docs/protocol/index.md @@ -40,35 +40,67 @@ msgpack maps with a `message_type` field. Arrays travel as `dtype` is a numpy typestr: a byte-order character, a kind character, and an item size. Parsing it takes about ten lines in any language. -## Four rules a first implementation usually gets wrong +## Rules a first implementation usually gets wrong **Messages strictly alternate.** `infer`, `action`, `feedback`, `infer`, and so on. The server's connection handler is straight-line code with no dispatcher, so a client that sends two `infer` messages in a row has the second one parsed as a `feedback` and is disconnected. -**The environment sets in one cycle need not match.** `infer` carries the -environments whose action chunk has run out; `feedback` carries the ones +**The action array is time-major.** `action` carries `[H, n, *da]`: the +horizon first, then the `n` environments of the `infer` it answers, in the +same order. The server lays it out environment-first internally and +transposes on the way out. A client that reads it environment-first runs +the wrong actions. It may use any prefix of the `H` steps and ask again +early, but never more than `H`. + +**`feedback`'s environment set need not match `infer`'s.** `infer` carries +the environments whose action chunk has run out; `feedback` carries the ones whose chunk finished on this step. The first time an environment terminates -early, those stop being the same set — permanently. The pairing between an -`action` and the `feedback` after it is flow control, not association; the -server routes feedback by environment index. +early, those stop being the same set — permanently. (The `action` always +answers exactly the `infer`'s set.) The pairing between an `action` and the +`feedback` after it is flow control, not association; the server routes +feedback by environment index. **The reward is the sum over the chunk.** Not the last step's. A client that reports the final step's reward trains a different MDP, and nothing fails. +**A done step reports its own observation.** On the step that sets +`terminated` or `truncated`, the observation in `feedback` must be the one +that step returned, not the first observation of the next episode. That is +Gymnasium's `AutoresetMode.NEXT_STEP`. An environment that resets inside its +own step sends the wrong one, and again nothing fails. + +**`info` is read for one key.** The servers read `info["episode"]` = +`{r, l, s, mask}` (return, length, success, is-this-a-finished-episode), each +an array of length `m`, the number of environments in that `feedback`. They +read it only on a transition whose `terminated` or `truncated` is set, and +feed it to `rollout/reward`, `rollout/length` and `rollout/success`. A +missing `mask` counts as true. Leaving `episode` out trains exactly the +same, but those three metrics read 0; that is how +[E44](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e44-cpp-pendulum)'s +C++ client found it. Nothing else in `info` is read, and `{}` is valid. +One trap, a Gap in SPEC.md: the server splits a non-empty `info` per +environment using the first length-`m` array it finds, at the top level or +one level into a nested map. With `m` > 1 and no such array - `{"task": +"pick"}`, or a msgpack list where an array belongs - the split comes out +wrong. The server treats that as a protocol error: it closes the connection +with the reason `plugrl-server-resync`, and its other clients go on. Until +plugrl-server #108 it closed with 1011 `Internal server error.` and shut the +whole server down. + **A reconnect starts from nothing.** Everything the server knows about an environment — its previous observation, the policy step state, its done flags — lives for exactly one connection. A client that reconnects must drop any `feedback` it was holding: the transition it describes can no longer be -completed, and sending it puts a transition built from an empty observation -into the training buffer. Nothing on either side reports an error when that -happens, which is what makes it worth stating. +completed. `plugrl-server` now logs `Feedback for env arrived with no +step state` and discards such a transition; before, it stored one built from +an empty observation without a word. The client still sees no error. ## Checking an implementation `plugrl-protocol` ships a server that grades a client against the -specification clause by clause and exits non-zero on a violation. +specification and exits non-zero on a violation. It is one command, and it runs your client for you: @@ -98,7 +130,7 @@ uv run --extra conformance plugrl-conformance --port 8000 --steps 20 \ `plugrl_client.cpp` uses POSIX sockets (`sys/socket.h`, `arpa/inet.h`), so that build needs Linux, macOS or WSL. The Python reference client has no such -constraint and exercises the same clauses: +constraint and goes through the same checker: ```bash uv run --extra conformance plugrl-conformance --port 8000 --steps 20 \ @@ -109,6 +141,22 @@ Its report has two severities. A **violation** is something the real server would reject or mishandle. A **note** is something it accepts that differs from what the Python client does — a portability risk, not a breach. +The checker watches one well-behaved connection, so it checks what that +connection shows: framing, alternation, environment indices, observation +shape, and the `feedback` payload's keys, dtypes and lengths. It does not +check the rest of the SPEC.md §8 checklist, and a clause it did not exercise +leaves no trace in its report. A client that breaks all of these still +prints "no violations": + +- the connection options (compression off, no frame size cap); +- reading `metadata` before sending anything; +- the chunk-summed reward, and the terminal observation on a done step; +- handling the close reasons `plugrl-server-stop` and `plugrl-server-resync` + differently, dropping held `feedback` on a reconnect, and treating a text + frame as fatal; +- what the client does with the `action` it receives: the time-major layout + and reading `env_ids`. + Two reference clients pass it. `raw_client.py` is 275 lines of Python using only `msgpack` and `websockets` — no numpy, nothing from PlugRL. `plugrl_client.cpp` is C++17 with **no third-party libraries at all**: SHA-1, @@ -116,8 +164,11 @@ base64, WebSocket framing and the msgpack subset the protocol needs are all in the one file, because that is the situation an embedded controller is actually in. -Both run in CI on every change, so the claim on this page stays checked -rather than remembered. +Both go through the checker in `plugrl-protocol`'s CI on every push to `main` +and every pull request. CI also sends the C++ client float64, time-major +actions and checks from its printed output that it decoded them. That is a +check of the C++ client, not something the checker can do for yours. Apart +from it, nothing in the list above is checked for either client. ## Relationship to openpi diff --git a/docs/protocol/index.zh.md b/docs/protocol/index.zh.md index 132fab0..2f53bd9 100644 --- a/docs/protocol/index.zh.md +++ b/docs/protocol/index.zh.md @@ -35,28 +35,53 @@ PlugRL 把一次训练拆成两个进程。**训练服务端**持有策略与学 `dtype` 是 numpy 的 typestr:一个字节序字符、一个类型字符、一个元素字节数。 用任何语言解析它大约十行代码。 -## 四条最容易实现错的规则 +## 最容易实现错的几条规则 **消息严格交替。** `infer`、`action`、`feedback`、`infer`……服务端的连接处理 是一段没有分发器的顺序代码,所以连发两个 `infer` 的客户端会让第二个被当成 `feedback` 解析,然后被断开。 -**一个周期里三条消息的环境集合不必相同。** `infer` 携带的是动作块刚用完的 +**动作数组是时间优先的。** `action` 的形状是 `[H, n, *da]`:先是 horizon,再是 +它所回应的那条 `infer` 里的 `n` 个环境,顺序不变。服务端内部按环境优先排,发出 +之前转置。按环境优先去读的客户端,执行的就是错的动作。客户端可以只用 `H` 步里 +的任意前缀、提前再要一次,但不能要超过 `H` 步。 + +**`feedback` 的环境集合不必与 `infer` 的相同。** `infer` 携带的是动作块刚用完的 环境,`feedback` 携带的是本步结束时动作块用完的环境。只要有一个环境提前终止, -这两个集合就**永久**不再相等。`action` 与其后 `feedback` 的配对只是流控, -不是语义关联 —— 服务端按环境编号查表路由。 +这两个集合就**永久**不再相等。(`action` 回应的永远正好是 `infer` 那一组。) +`action` 与其后 `feedback` 的配对只是流控,不是语义关联 —— 服务端按环境编号查表路由。 **奖励是整个动作块上的求和**,不是最后一步的奖励。只汇报最后一步的客户端会在 一个不同的 MDP 上训练,而且不会有任何东西报错。 +**结束的那一步汇报它自己的观测。** 在置了 `terminated` 或 `truncated` 的那一步, +`feedback` 里的观测必须是这一步返回的那一帧,而不是下一个 episode 的第一帧。 +这就是 Gymnasium 的 `AutoresetMode.NEXT_STEP`。在自己的 step 里就 reset 的环境 +会发错这一帧,同样不会有任何报错。 + +**`info` 只读一个键。** 服务端只读 `info["episode"]` = `{r, l, s, mask}`(回报、 +长度、是否成功、这一项是否是刚结束的 episode),每一项都是长度为 `m` 的数组, +`m` 是这条 `feedback` 里的环境数。只有在置了 `terminated` 或 `truncated` 的转移 +上才读,读来的值进 `rollout/reward`、`rollout/length` 和 `rollout/success`。缺了 +`mask` 就当作 true。不发 `episode`,训练完全一样,只是这三个指标一直是 0; +[E44](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e44-cpp-pendulum) +的 C++ 客户端就是这么发现它的。`info` 里的其他内容一概不读,发 `{}` 也合法。 +有一个坑,是 SPEC.md 里的一条 Gap:服务端拆分非空的 `info` 时,靠的是它找到的第一个 +长度为 `m` 的数组,只在顶层或往下一层的嵌套 map 里找。`m` > 1 而又没有这样的数组时 +—— 比如 `{"task": "pick"}`,或者该放数组的地方放了 msgpack list —— 就会拆错。 +服务端把这当作协议错误:以 `plugrl-server-resync` 为原因关闭这条连接,其他客户端 +照常继续。在 plugrl-server #108 之前,连接会以 1011 `Internal server error.` 关闭, +整个服务端也随之退出。 + **重连意味着从零开始。** 服务端关于一个环境的全部记忆 —— 上一帧观测、策略的 step state、终止标志 —— 只活在一条连接里。重连的客户端必须丢弃手上未发出的 -`feedback`:它描述的那次转移已经无法补全,发出去只会往训练缓冲里塞一条由空 -观测拼出来的转移。这件事发生时两端都不会报错,所以才值得单独写一条。 +`feedback`:它描述的那次转移已经无法补全。`plugrl-server` 现在遇到这种转移会打一条 +`Feedback for env arrived with no step state` 警告并把它丢掉;以前它会不声不响 +地存下一条由空观测拼出来的转移。客户端这边仍然看不到任何报错。 ## 检验一个实现 -`plugrl-protocol` 附带一个服务端,它按规范逐条给客户端打分,有违规就以非零码退出。 +`plugrl-protocol` 附带一个服务端,它按规范给客户端打分,有违规就以非零码退出。 一条命令就够,它会替你把客户端也起起来: @@ -83,7 +108,7 @@ uv run --extra conformance plugrl-conformance --port 8000 --steps 20 \ `plugrl_client.cpp` 用的是 POSIX socket(`sys/socket.h`、`arpa/inet.h`), 所以这一步需要 Linux、macOS 或 WSL。Python 参考客户端没有这个限制, -覆盖的条款是同一批: +过的是同一个检验器: ```bash uv run --extra conformance plugrl-conformance --port 8000 --steps 20 \ @@ -93,12 +118,27 @@ uv run --extra conformance plugrl-conformance --port 8000 --steps 20 \ 报告分两个等级。**violation** 是真服务端会拒绝或处理错的问题;**note** 是真服务端 接受、但与 Python 客户端做法不同的地方 —— 是可移植性风险,不是违约。 +检验器看的是一条规规矩矩的连接,所以它只查这条连接看得到的东西:分帧、消息交替、 +环境编号、观测形状,以及 `feedback` 载荷的键、dtype 和长度。SPEC.md §8 清单里的其余 +条款它不查,而且没被触及的条款在报告里不留任何痕迹。下面这些全违反的客户端,照样 +打印 "no violations": + +- 连接选项(关闭压缩、不限帧大小); +- 发任何东西之前先读 `metadata`; +- 按动作块求和的奖励,以及结束那一步的终止观测; +- 区别对待 `plugrl-server-stop` 和 `plugrl-server-resync` 两种关闭原因、重连时丢弃 + 手上的 `feedback`、把文本帧当致命错误; +- 客户端拿到 `action` 之后怎么用:时间优先的布局,以及读 `env_ids`。 + 两个参考客户端都能通过。`raw_client.py` 是 275 行 Python,只用 `msgpack` 和 `websockets`,不用 numpy,也不用 PlugRL 的任何东西。`plugrl_client.cpp` 是 C++17, **完全不依赖第三方库**:SHA-1、base64、WebSocket 分帧,以及协议需要的那部分 msgpack,全都写在同一个文件里 —— 因为嵌入式控制器面对的就是这种处境。 -两者都在每次改动的 CI 里跑,所以本页的说法是被持续检验的,而不是被记住的。 +`plugrl-protocol` 的 CI 在每次推到 `main` 和每个 pull request 上都让两者过一遍检验器。 +CI 还会给 C++ 客户端发 float64、时间优先的动作,再从它打印的输出核对它解对了。 +这是对 C++ 客户端的检查,检验器没法替你的客户端做。除此之外,上面清单里的各条对 +两个客户端都没有被检查。 ## 与 openpi 的关系 diff --git a/docs/user_guide/get_started.md b/docs/user_guide/get_started.md index 6e83d8f..81d0c33 100644 --- a/docs/user_guide/get_started.md +++ b/docs/user_guide/get_started.md @@ -14,27 +14,33 @@ cd plugrl-server && uv sync && cd .. cd plugrl-env-client && uv sync --extra mujoco && cd .. ``` +`uv sync` puts each package, and its `plugrl-run-*` command, in that +repository's own `.venv`. So every command on these pages is run from inside +the repository it belongs to, with `uv run`, which uses that `.venv`. + `--extra mujoco` is what the quickstart below needs. The env client has one -extra per environment family; install only the ones you use. +extra per environment family; install only the ones you use. `uv sync` +removes whatever it was not asked for, so pass the same `--extra` flags every +time you run it. `uv run` leaves them alone. -They can also live in one environment - the two dependency sets do coexist, -which is measured in `plugrl-server/experiments/e1-dependency-conflict/`. +The two packages can also live in one environment - the two dependency sets +do coexist, which is measured in `plugrl-server/experiments/e1-dependency-conflict/`. Separate environments are simply the point of the split. ## Quickstart -Terminal A, the training server: +Terminal A, the training server, in `plugrl-server`: ```bash -plugrl-run-server fpo-policy default fpo default \ +uv run plugrl-run-server fpo-policy default fpo default \ --port 8000 --policy.device cpu \ --algo.global-steps 500000 --algo.buffer-size 4096 ``` -Terminal B, the environment: +Terminal B, the environment, in `plugrl-env-client`: ```bash -plugrl-run-env-client mujoco-v1 \ +uv run plugrl-run-env-client mujoco-v1 \ --server-host 127.0.0.1 --server-port 8000 \ --num-envs 1 --num-episodes 600 --runner.replan-steps 1 --runner.seed 0 ``` @@ -52,10 +58,14 @@ are `fpo-policy`'s defaults, so nothing needs configuring. shorter than about a million steps learns **once, at the very end**, giving a single point instead of a curve. +`--server-host 127.0.0.1` is not optional either. The server listens on +`0.0.0.0`, meaning every interface. The client's default `--server-host` is +also `0.0.0.0`, and on Windows that is not an address a client can connect to. + ## Verify -- The server prints a WebSocket listening address. -- The env client prints the server's metadata, including `action_dim` and +- The server prints `Agent Server is listening on 0.0.0.0:8000`. +- The env client prints `Server metadata: {...}`, including `action_dim` and `action_horizon`, and starts stepping episodes. - The server's metrics show `rollout/reward` rising. On HalfCheetah it starts near -300 and climbs out within a few minutes. @@ -63,24 +73,48 @@ are `fpo-policy`'s defaults, so nothing needs configuring. `plugrl-server/experiments/e6-first-learning-curve/` holds a three-seed run of exactly this, with the script that produced it. +## Where output goes + +- Server: `////`, for + example `checkpoints/fpo/fpo-policy//`. It holds one directory per + saved step and a `tensorboard/` directory with the metrics. + `--checkpoint-base-dir` defaults to `./checkpoints`. Without `--exp-name` + the server makes up a name from the time. +- Env client: `runs//`, under the directory the client was started + from. `client_config.json` has every setting the client ran with, + `logs/client.log` its log, and `rollout/proc_000/summary.json` the episode + count, mean return, success rate and a timing breakdown. + ## Just checking connectivity ```bash -plugrl-run-server dummy-policy default dummy default -plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server-port 8000 +# Terminal A, in plugrl-server +uv run plugrl-run-server dummy-policy default dummy default + +# Terminal B, in plugrl-env-client +uv run plugrl-run-env-client dummy-v1 --server-host 127.0.0.1 --server-port 8000 --num-episodes 3 ``` -The dummy algorithm's `learn` is a sleep and moves no weights. Use it to -confirm the two sides talk, not to train. +The dummy policy's default action - continuous, 7-dimensional, horizon 4 - +is what `dummy-v1` expects, so neither side needs more flags. The client +exits 0 after three episodes. The dummy algorithm's `learn` is a sleep and +moves no weights. Use it to confirm the two sides talk, not to train. ## Other policies ```bash -plugrl-run-server dppo-policy default dppo hopper --exp_name my_dppo_exp +uv run plugrl-run-server dppo-policy default dppo hopper \ + --policy.checkpoint-path /path/to/pretrained.pt --exp-name my_dppo_exp ``` -DPPO needs `plugrl-server[dppo]` and a pretrained checkpoint. `pi0-policy` -needs a checkpoint too, and a GPU. +`dppo-policy` needs `uv sync --extra dppo`; without it the policy is not in +the CLI at all. It also needs a pretrained checkpoint: without +`--policy.checkpoint-path` it starts from random weights. + +`pi0-policy` needs a GPU, a checkpoint, the `openpi` extra and the +`third_party/openpi` git submodule, which `git clone` does not fetch. Until +all of that is installed it is missing from the CLI. The steps are in the +[plugrl-server README](https://github.com/PlugRL/plugrl-server#training-pi0-openpi-with-fpo). !!! note "The Ray launcher is not a supported path today" @@ -95,17 +129,21 @@ needs a checkpoint too, and a GPU. - Set the server address with `--host` and `--port`. - Control the episode count with `--num-episodes`. - `--num-procs` runs several env client processes against one server. -- `--resume` requires an existing experiment directory under - `--checkpoint-base-dir`. +- `--resume` continues a run. Pass the same `--exp-name`, policy and + algorithm as the run you are continuing (and the same + `--checkpoint-base-dir`, if you set one), because together they name the + checkpoint directory. Without `--exp-name` the server makes up a new name, + finds no checkpoint there, and raises `FileNotFoundError`. ## Troubleshooting | Symptom | Cause | |---|---| -| Env client keeps retrying | The server is not listening yet, or the address or firewall is wrong | +| Env client keeps retrying | The server is not listening yet, the address or firewall is wrong, or `--server-host` was left at `0.0.0.0` | | `Torch not compiled with CUDA enabled` | Pass `--policy.device cpu` | | Only one metrics row, at the very end | `--algo.buffer-size` is larger than the run | -| `--resume` raises `FileNotFoundError` | There are no checkpoints in that directory yet | +| `--resume` raises `FileNotFoundError` | No checkpoint in `////`: `--exp-name`, policy or algorithm differs from the run you meant, or that run saved nothing yet | +| `FileExistsError: Checkpoint directory ... already exists` | That `--exp-name` was used before. Pass `--resume`, `--overwrite` (deletes it), or a new name | | An environment reports a missing extra | Install it, e.g. `uv sync --extra mujoco` | ## Next steps diff --git a/docs/user_guide/get_started.zh.md b/docs/user_guide/get_started.zh.md index 38c93ee..6839f34 100644 --- a/docs/user_guide/get_started.zh.md +++ b/docs/user_guide/get_started.zh.md @@ -4,7 +4,7 @@ ## 安装 -两个包都不在 PyPI 上。分别 clone 并用 `uv` 安装: +两个包都不在 PyPI 上。分别 clone,再用 `uv` 安装: ```bash git clone https://github.com/PlugRL/plugrl-server.git @@ -14,33 +14,37 @@ cd plugrl-server && uv sync && cd .. cd plugrl-env-client && uv sync --extra mujoco && cd .. ``` -`--extra mujoco` 是下面快速开始所需的。env client 的每个环境家族对应一个 extra, -只装你要用的即可。 +`uv sync` 把每个包连同它的 `plugrl-run-*` 命令装进各自仓库的 `.venv`。所以这些 +页面上的命令都要在它所属的仓库目录里、用 `uv run` 来跑,`uv run` 用的就是那个 `.venv`。 -两者也可以装在同一个环境里 —— 这两套依赖确实能共存, +`--extra mujoco` 是下面快速开始要用的。env client 每个环境家族对应一个 extra, +用哪个装哪个。`uv sync` 会删掉这次没点名的东西,所以每次跑它都要带上同样的 +`--extra`;`uv run` 不会动它们。 + +两个包也可以装在同一个环境里 —— 这两套依赖确实能共存, `plugrl-server/experiments/e1-dependency-conflict/` 里有实测。 分开装只是这个架构的本意。 ## 快速开始 -终端 A,训练端: +终端 A,训练端,在 `plugrl-server` 里: ```bash -plugrl-run-server fpo-policy default fpo default \ +uv run plugrl-run-server fpo-policy default fpo default \ --port 8000 --policy.device cpu \ --algo.global-steps 500000 --algo.buffer-size 4096 ``` -终端 B,环境端: +终端 B,环境端,在 `plugrl-env-client` 里: ```bash -plugrl-run-env-client mujoco-v1 \ +uv run plugrl-run-env-client mujoco-v1 \ --server-host 127.0.0.1 --server-port 8000 \ --num-envs 1 --num-episodes 600 --runner.replan-steps 1 --runner.seed 0 ``` -`HalfCheetah-v5` 是 17 维观测、6 维动作,**正好是 `fpo-policy` 的默认值**, -所以不需要任何配置。 +`HalfCheetah-v5` 是 17 维观测、6 维动作,正好是 `fpo-policy` 的默认值, +所以什么都不用配。 !!! warning "两个不是装饰的参数" @@ -51,34 +55,57 @@ plugrl-run-env-client mujoco-v1 \ 一步时才学习。按默认的 `983040`,少于约一百万步的运行**只会在最后学一次**, 你得到的是一个点而不是一条曲线。 +`--server-host 127.0.0.1` 同样不能省。server 监听的是 `0.0.0.0`,即所有网卡; +客户端 `--server-host` 的默认值也是 `0.0.0.0`,而在 Windows 上客户端连不上这个地址。 + ## 验证 -- server 打印 WebSocket 监听地址 -- env client 打印 server 元信息(含 `action_dim`、`action_horizon`)并开始跑 episode +- server 打印 `Agent Server is listening on 0.0.0.0:8000` +- env client 打印 `Server metadata: {...}`(含 `action_dim`、`action_horizon`), + 然后开始跑 episode - server 的指标里 `rollout/reward` 在上升。HalfCheetah 上从 -300 附近起步, 几分钟内就会爬上来 `plugrl-server/experiments/e6-first-learning-curve/` 里有一次三种子的完整运行, 以及产生它的脚本。 +## 输出在哪里 + +- server:`////`,例如 + `checkpoints/fpo/fpo-policy//`。里面每保存一步就有一个目录,另有一个 + 放指标的 `tensorboard/` 目录。`--checkpoint-base-dir` 默认是 `./checkpoints`; + 不传 `--exp-name` 时,server 用当前时间生成一个名字。 +- env client:启动目录下的 `runs//`。`client_config.json` 记着这次运行 + 的全部设置,`logs/client.log` 是日志,`rollout/proc_000/summary.json` 里有 + episode 数、平均回报、成功率和耗时拆分。 + ## 只想确认能连通 ```bash -plugrl-run-server dummy-policy default dummy default -plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server-port 8000 +# Terminal A, in plugrl-server +uv run plugrl-run-server dummy-policy default dummy default + +# Terminal B, in plugrl-env-client +uv run plugrl-run-env-client dummy-v1 --server-host 127.0.0.1 --server-port 8000 --num-episodes 3 ``` -dummy 算法的 `learn` 是一个 sleep,不移动任何权重。它用来确认两端能对话, -不是用来训练的。 +dummy 策略默认输出连续的 7 维动作、horizon 为 4,正是 `dummy-v1` 要的,所以两边 +都不用再加参数。客户端跑完三个 episode 后以 0 退出。dummy 算法的 `learn` 只是 +sleep,不更新任何权重。它用来确认两端能对话,不是用来训练的。 ## 其他策略 ```bash -plugrl-run-server dppo-policy default dppo hopper --exp_name my_dppo_exp +uv run plugrl-run-server dppo-policy default dppo hopper \ + --policy.checkpoint-path /path/to/pretrained.pt --exp-name my_dppo_exp ``` -DPPO 需要 `plugrl-server[dppo]` 和一个预训练 checkpoint。`pi0-policy` 同样需要 -checkpoint,而且需要 GPU。 +`dppo-policy` 需要 `uv sync --extra dppo`,不装的话 CLI 里根本没有它。它还需要 +一个预训练 checkpoint:不传 `--policy.checkpoint-path` 就是从随机权重开始。 + +`pi0-policy` 需要 GPU、checkpoint、`openpi` extra,以及 `git clone` 不会拉下来的 +`third_party/openpi` git 子模块。这些没装齐之前,CLI 里不会出现它。步骤见 +[plugrl-server 的 README](https://github.com/PlugRL/plugrl-server#training-pi0-openpi-with-fpo)。 !!! note "Ray 启动器目前不是受支持的路径" @@ -92,16 +119,20 @@ checkpoint,而且需要 GPU。 - 用 `--host` 和 `--port` 设置服务端地址 - 用 `--num-episodes` 控制 episode 数 - `--num-procs` 可以起多个环境端进程连同一个 server -- `--resume` 需要 `--checkpoint-base-dir` 下已有实验目录 +- `--resume` 接着之前的运行继续。`--exp-name`、策略、算法都要和那次运行一样 + (设过 `--checkpoint-base-dir` 的话也要一样),因为 checkpoint 目录就是由它们 + 拼出来的。不传 `--exp-name` 时 server 会新起一个名字,在那里找不到 checkpoint, + 于是报 `FileNotFoundError`。 ## 排错 | 现象 | 原因 | |---|---| -| env client 一直重试 | server 还没监听,或地址/防火墙不对 | +| env client 一直重试 | server 还没监听,地址或防火墙不对,或者 `--server-host` 还是默认的 `0.0.0.0` | | `Torch not compiled with CUDA enabled` | 加上 `--policy.device cpu` | | 指标只有最后一行 | `--algo.buffer-size` 比整个运行还大 | -| `--resume` 报 `FileNotFoundError` | 那个目录下还没有任何 checkpoint | +| `--resume` 报 `FileNotFoundError` | `////` 下没有 checkpoint:`--exp-name`、策略或算法和你想接的那次不一样,或者那次还没存过 | +| `FileExistsError: Checkpoint directory ... already exists` | 这个 `--exp-name` 用过了。加 `--resume`、`--overwrite`(会删掉旧目录),或者换个名字 | | 某个环境报缺少 extra | 装上它,例如 `uv sync --extra mujoco` | ## 下一步 diff --git a/docs/user_guide/index.md b/docs/user_guide/index.md index 0549ac8..488cd68 100644 --- a/docs/user_guide/index.md +++ b/docs/user_guide/index.md @@ -4,43 +4,57 @@ Run PlugRL end to end: start a server, then start one or more env clients. ## Quickstart +After the install in [Get Started](get_started.md), run each command from +inside its own repository, in its own terminal: + ```bash -plugrl-run-server fpo-policy default fpo default \ +# in plugrl-server +uv run plugrl-run-server fpo-policy default fpo default \ --policy.device cpu --algo.global-steps 500000 --algo.buffer-size 4096 -plugrl-run-env-client mujoco-v1 --server-host 127.0.0.1 --server-port 8000 \ + +# in plugrl-env-client +uv run plugrl-run-env-client mujoco-v1 --server-host 127.0.0.1 --server-port 8000 \ --num-envs 1 --num-episodes 600 --runner.replan-steps 1 --runner.seed 0 ``` -That pair learns. [Get Started](get_started.md) explains the two flags that -are not optional, and has the dummy connectivity check. +That pair learns. [Get Started](get_started.md) explains the two server flags +that are not optional, and has the dummy connectivity check. ## Verify -- Server prints a WebSocket listening address. -- Env client prints server metadata and steps episodes. +- The server prints `Agent Server is listening on 0.0.0.0:8000`. +- The env client prints `Server metadata: {...}` and steps episodes. ## Workflow -1. Start a training server with `plugrl-run-server`. (There is also - `plugrl-run-server-ray`, but it is not a supported path today - see - [Get Started](get_started.md).) -2. Start one or more env clients with `plugrl-run-env-client `. +1. Start a training server with `uv run plugrl-run-server` in `plugrl-server`. + (There is also `plugrl-run-server-ray`, but it is not a supported path + today - see [Get Started](get_started.md).) +2. Start one or more env clients with `uv run plugrl-run-env-client ` + in `plugrl-env-client`. ## Components -- `plugrl-server`: batches inference across connected workers, runs learning and checkpointing +- `plugrl-server`: batches inference across connected clients, runs learning and checkpointing - `plugrl-env-client`: creates Gymnasium envs, sends `infer`, receives `action`, sends `feedback` - `plugrl-protocol`: WebSocket transport, message types, and msgpack serialization ## Common options -- Server default address is `0.0.0.0:8000`. -- Env client connects via `--server-host` and `--server-port`. +- The server listens on `0.0.0.0:8000` by default, which means every + interface. Set it with `--host` and `--port`. +- The env client connects to `--server-host` and `--server-port`. Always pass + `--server-host`: its default is `0.0.0.0`, which a client on Windows cannot + connect to. On the server's own machine use `127.0.0.1`. ## Troubleshooting -- Env client keeps retrying: confirm the server is listening and the address is reachable. -- Policy or algorithm not found in CLI: ensure the registration module is imported before the CLI is built. +- Env client keeps retrying: the server is not listening yet, or + `--server-host` / `--server-port` is wrong. +- An env ID or policy is missing from the CLI: its extra is not installed, or + (for your own code) its module was never imported. The env client prints a + `Skip loading env module ...` warning that names the missing extra; the + server leaves the policy out without a message. ## Next steps diff --git a/docs/user_guide/index.zh.md b/docs/user_guide/index.zh.md index 2cf2506..afb3730 100644 --- a/docs/user_guide/index.zh.md +++ b/docs/user_guide/index.zh.md @@ -1,45 +1,55 @@ # 用户指南 -端到端跑通 PlugRL:启动 server,启动 env client。 +端到端跑通 PlugRL:先启动 server,再启动一个或多个 env client。 ## 快速开始 +按[快速开始](get_started.zh.md)装好之后,每条命令都在它所属的仓库目录里、 +各开一个终端运行: + ```bash -plugrl-run-server fpo-policy default fpo default \ +# in plugrl-server +uv run plugrl-run-server fpo-policy default fpo default \ --policy.device cpu --algo.global-steps 500000 --algo.buffer-size 4096 -plugrl-run-env-client mujoco-v1 --server-host 127.0.0.1 --server-port 8000 \ + +# in plugrl-env-client +uv run plugrl-run-env-client mujoco-v1 --server-host 127.0.0.1 --server-port 8000 \ --num-envs 1 --num-episodes 600 --runner.replan-steps 1 --runner.seed 0 ``` -这一对**真的会学**。[快速开始](get_started.zh.md)里说明了那两个不可省的参数, +这一对真的会学。[快速开始](get_started.zh.md)里说明了 server 那两个不能省的参数, 以及 dummy 连通性检查怎么做。 ## 验证 -- server 打印 WebSocket 监听地址 -- env client 打印 server 元信息并开始跑 episode +- server 打印 `Agent Server is listening on 0.0.0.0:8000` +- env client 打印 `Server metadata: {...}`,然后开始跑 episode ## 流程 -1. 用 `plugrl-run-server` 启动训练端。(也有 `plugrl-run-server-ray`,但它目前 - 不是受支持的路径,见[快速开始](get_started.zh.md)) -2. 用 `plugrl-run-env-client ` 启动一个或多个环境端。 +1. 在 `plugrl-server` 里用 `uv run plugrl-run-server` 启动训练端。(另有 + `plugrl-run-server-ray`,但它目前不是受支持的路径,见[快速开始](get_started.zh.md)) +2. 在 `plugrl-env-client` 里用 `uv run plugrl-run-env-client ` 启动一个或多个环境端。 ## 组件 -- `plugrl-server`:聚合推理请求,驱动学习与 checkpoint +- `plugrl-server`:把各个客户端的推理请求合批,负责学习与 checkpoint - `plugrl-env-client`:创建 Gymnasium 环境,发送 `infer`,接收 `action`,回传 `feedback` -- `plugrl-protocol`:WebSocket 传输与 msgpack 序列化 +- `plugrl-protocol`:WebSocket 传输、消息类型与 msgpack 序列化 ## 常用参数 -- server 默认地址为 `0.0.0.0:8000` -- env client 通过 `--server-host` 与 `--server-port` 连接 +- server 默认监听 `0.0.0.0:8000`,即所有网卡。用 `--host`、`--port` 修改。 +- env client 连接 `--server-host` 与 `--server-port`。`--server-host` 一定要传: + 它的默认值是 `0.0.0.0`,Windows 上的客户端连不上这个地址。和 server 在同一台 + 机器上就用 `127.0.0.1`。 ## 常见问题 -- env client 一直重试:确认 server 已监听且地址可达。 -- CLI 里找不到策略或算法:确认注册模块在构建 CLI 前已被 import。 +- env client 一直重试:server 还没开始监听,或者 `--server-host` / `--server-port` 写错了。 +- CLI 里找不到某个 env ID 或策略:对应的 extra 没装,或者(你自己写的代码)模块 + 根本没被 import。env client 启动时会打印一条 `Skip loading env module ...` 警告, + 里面写着缺哪个 extra;server 则不声不响地把这个策略略过。 ## 下一步