Skip to content

Add reasoning-guided pi0.5 policy-learning pilot - #705

Draft
rpuns wants to merge 37 commits into
astra/demo-skill-library-20260930from
astra/reasoning-policy-learning-20261003
Draft

rpuns wants to merge 37 commits into
astra/demo-skill-library-20260930from
astra/reasoning-policy-learning-20261003

Conversation

@rpuns

@rpuns rpuns commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

This study tests whether Astra can teach autonomous OOD behavior more sample-efficiently than RL from the same π0.5 checkpoint. Astra observes real cameras, state and execution history, compares computational corrections with one fixed native proposal, executes one selected prefix, and updates policy LoRA parameters from useful real action windows. FRS action steering, physical candidate retries, privileged object poses and synthetic training targets are excluded.

The research objective remains unmet. No completed checkpoint reaches the preregistered 8/10 autonomous target, and no fresh-task or untouched-reset confirmation was performed. Measured scope is Goal OOD6 / seed173 and Spatial OOD2 / seeds173 and179. These development runs are separate from the earlier 20-task intervention campaign. Each SR uses ten declared reset states with simulator success and verified reset identities.

Measured condition Native Learned Collection controls
V8, Goal seed173, four collections 5/10 5/10 815
V7, Spatial seed173, two collections 3/10 5/10 620
V7, Spatial seed179, two collections 2/10 2/10 416
Local labels from one native success, Spatial seed179 2/10 3/10 106 reused
Whole-success BC, same source and seed 2/10 2/10 106 reused
Tuned DSRL, Goal seed173, eight collections 5/10 6/10 1,455
Tuned PPO, Goal seed173, eight collections 5/10 4/10 1,686
Tuned DSRL, Spatial seed173, eight collections 3/10 3/10 2,299
Tuned PPO, Spatial seed173, eight collections 3/10 2/10 2,480
Tuned DSRL, Spatial seed179, two collections 2/10 2/10 411

These endpoints have different budgets; the dashboard retains all 64 completed scheduled curve points with their actual interaction and token costs. The first Spatial teacher gain has not repeated. Its later policy3 was saved but remained unevaluated after external quota preemption.

The final paired control freezes one unassisted native success before either student outcome. Fourteen locally useful windows yield 3/10, versus 2/10 from all eighteen successful-episode windows. Both fresh students receive 40 uniform seven-channel updates, with identical initial adapter hashes and sampled flow-time sequences. Only the local-label student gains reset20; neither loses a native win. This is one extra successful reset, not confirmed superiority. Both branches retain the original 106 controls and 455,499 monitoring tokens; they add no collection or teacher calls and 5,168 evaluation controls. This is not the original online learning trajectory or a measured teacher-free collection.

The intervention audit exposes a bottleneck: the seed179 teacher executes ten locally useful assisted commands but none enters its complete-window training data. Relaxing selection is insufficient. On the same Goal data, local labels / hindsight / whole-success BC score 5/10 / 4/10 / 5/10 at 100 updates. Strict replay scores 5/10, 5/10, 4/10 at 40/100/200 updates; evidence-masked replay scores 5/10, 3/10, 0/10 despite improving training fit. Versioned revisions include RTC weighting, gripper targets, prompt/TEI/TLI candidates, bounded controller targets and per-rollout visual grounding; recorded decisions distinguish offered methods from actual execution.

The frozen V8 evaluation restores the exact saved policy4 and repeats all ten original resets without new training or teacher calls. It preserves the native binary outcomes and adds 1,951 evaluation controls. All seven earlier completed outcomes reproduce; six command sequences match exactly. One failed trajectory diverges after initially identical commands. Earlier retained evaluation costs remain charged; totals remain lower bounds wherever an unarchived tail is unknown. Repeated cases do not become a seventeen-trial denominator.

The study closes at 22.055960 of 24 authorized L40S GPU-hours, including initialization and failed workers. All 26 GPU allocations have ended, with peak concurrency two. Retained totals are at least 18,101 collection controls, 115,777 evaluation controls and 38,616,918 reported teacher tokens including cache and offline diagnostics. Reused source costs count once globally. Tokens are reported usage, not an API bill.

Reproducibility includes pinned full weights and RLinf components, strict native/converted-weight checks, immutable OSMO bundles, exact Astra prompts, saved checkpoints, per-reset outcomes, checksummed execution/judgment/admission/optimizer records, actual camera examples, rollout videos and a standalone HTML dashboard with CSV/JSON/PNG/PDF exports. The serial RL integrations have documented overrides and limited tuning; this is not a stock SOTA benchmark reproduction. DSRL's initial noise distribution differs from the native Gaussian despite sharing decoder weights.

Validation: full pytest tests/unit — 1,794 passed, six optional skips. Changed Python passes Ruff. Chrome-rendered tables, JavaScript syntax, archive integrity, record checksums and privacy scans pass. Interrupted evaluation accounting has a regression test; private provider event streams and signed catalogs are excluded from public record export.

Linear stack: based on astra/demo-skill-library-20260930 (PR #704), with no merges or rewritten history.

rpunamiya added 30 commits October 3, 2026 14:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant