Stop resetting the environment after the last episode - #2
Merged
Merged
Conversation
rollout() reset unconditionally whenever an episode finished, including the
one that satisfied num_episodes. The observation it produced was never read -
the loop exits immediately afterwards. For a simulator that is a wasted
rollout. For a real robot it is a pointless move back to the home pose after
the run is already over.
The test that proves it comes from nothingbutbut's features/robocasa branch,
where this was fixed and never merged. It is the only test in this package
that exercises rollout() end to end other than the two added yesterday, so
it is worth having on main whatever happens to the rest of that branch.
Two adaptations were needed, both because main has moved since April:
* its fake agent returned an action of shape (steps, n_envs), which was
only valid under the batch-axis guess that _get_action_spec made until
yesterday. It now returns [H, n, *da] like the protocol specifies.
* its recorder spy predates record_timing.
test_cli.py::test_rollout_infer_feedback_and_partial_reset_semantics had to
change with it, and this is worth being explicit about: it asserted that a
partial reset happens, using num_episodes=1 as the shortest route to a done
environment. Under the new behaviour that run is over before the reset would
happen, so the test now uses two episodes. Its subject - that a partial reset
targets only the environments that finished - is unchanged and still
asserted.
Co-Authored-By: nothingbutbut <2367347983@qq.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
rollout() reset unconditionally whenever an episode finished, including
the one that satisfied
num_episodes. The observation it produced wasnever read - the loop exits immediately afterwards. For a simulator that is
a wasted rollout; for a real robot it is a pointless move back to the home
pose after the run is over.
The test that proves it is nothingbutbut's, from the unmerged
features/robocasabranch, where this was fixed in April. It is the onlyend-to-end
rollout()test in the package other than the two addedyesterday, so it belongs on main whatever happens to the rest of that
branch.
Two adaptations, both because main has moved since April: the fake agent
returned
(steps, n_envs), valid only under the_get_action_specbatch-axis guess that was removed yesterday, and now returns
[H, n, *da];and the recorder spy predates
record_timing.test_rollout_infer_feedback_and_partial_reset_semanticschanged with it,which is worth being explicit about. It asserted that a partial reset
happens, using
num_episodes=1as the shortest route to a doneenvironment. Under the new behaviour that run is over before the reset would
happen, so it now uses two episodes. Its subject - a partial reset targets
only the environments that finished - is unchanged and still asserted.
🤖 Generated with Claude Code