Skip to content

E37: from the same clone, fpo-policy learns square under DPPO and falls under our FPO - #89

Merged
tactino merged 10 commits into
mainfrom
exp/e37-square-fpo-policy
Sep 28, 2026
Merged

tactino merged 10 commits into
mainfrom
exp/e37-square-fpo-policy

Conversation

@tactino

@tactino tactino commented Sep 28, 2026

Copy link
Copy Markdown
Member

E37 was pre-registered at fb93a91, after two pilots. It fine-tunes E35's behaviour-cloned fpo-policy on robomimic square under both algorithms: three seeds, three client processes each, on guangzhao.

cell start (seeds 0/1/2) end: last ten change
dppo-square (40 iterations) 0.390 / 0.411 / 0.372 0.781 / 0.784 / 0.798 +0.37 to +0.43
fpo-square (60 iterations, FPO++'s square settings as far as our FPO has them) 0.517 / 0.531 / 0.550 0.256 / 0.353 / 0.392 −0.16 to −0.26
  • P1 (runs) and V1 (settings took) hold.
  • P3 holds: fpo-policy · DPPO · square learns, 3 of 3.
  • P2 is falsified: under FPO it falls on every seed. It held for 20–40 iterations and then drifted down, the same shape as E36's FPO++ arms on pi0.5.

FINDINGS also records something read from the logs, not registered. FPO's critic never fits the return. Its value loss went only from about 9×10⁴ to about 4×10⁴ over 60 iterations, and the raw advantage std stayed near 400. DPPO's explained variance went from 0.40 to 0.61. The FPO arm gave its critic the actor's learning rate (1e-5, with rewards ×10). FPO++ gives its critic 1e-4. PROTOCOL.md declared that as a known deviation, and it is the first thing to try next. By this project's rule this is a defect of ours, not a finding about FPO.

The figure update (fpo-dppo-square → learns, fpo-fpo-square from E37's runs) will go into #88 and the site PR.

…-ratio' and 'origin/feat/dppo-square-variant' into exp/e37-square-fpo-policy
…ur-cloned start

E35's clone (0.44 / 0.38 success), observation statistics frozen (#82), the
critic from the run's own initialisation; fpo-square under FPO with FPO++'s
square fine-tuning (60 iterations of 48,000 steps), dppo-square under dppo
square with E33's flow-policy changes (40 of 80,000); three seeds each.
Learns: success over the last ten iterations at least +0.2 over the
iterations the clone collected unchanged, on 2 of 3 (E34's rule). Written
after the two pilots in results/pilot.txt, before the registered run.
…r DPPO (3 of 3) and falls under FPO (0 of 3)
@tactino
tactino merged commit de4b3db into main Sep 28, 2026
3 checks passed
@tactino
tactino deleted the exp/e37-square-fpo-policy branch September 28, 2026 19:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant