Skip to content

fix(evolve): label the swarm loop honestly as a demonstration - #24

Merged
ZuluYokohama merged 3 commits into
masterfrom
fix/honest-swarm-labeling
Jun 6, 2026
Merged

ZuluYokohama merged 3 commits into
masterfrom
fix/honest-swarm-labeling

Conversation

@ZuluYokohama

@ZuluYokohama ZuluYokohama commented Jun 6, 2026 •

Copy link
Copy Markdown
Collaborator

fix(evolve): label the "swarm" loop honestly as a demonstration

Branch-target note: Fork has no dev branch; default is master. Internal fork PR → master.

Who is submitting this PR? (required)

Field Value
Your model + version Claude Opus 4.8 (claude-opus-4-8) orchestration; opus-tier subagents (exact minor version not surfaced; likely 4.8).
Harness + version Claude Code (CLI); version not surfaced.
All plugins installed superpowers 5.1.0, plugin-dev, context7 (MCP).
Human partner who reviewed this diff b.jones@jtech.ai (ZuluYokohama)

What problem are you trying to solve?

scripts/evolve.py and scripts/recursive_evolution.py are described as an "Infinite Recursive LLM Swarm," and their output claims real work — "Dispatching parallel swarm agents…", "Swarm returned 3 potential intent paths", "Evolution successful! Feature tallied & mapped", "Branch cleanly merged", "achieved absolute MaxVal". In reality the loop performs no LLM work: recursive_evolution.py deterministically flips each feature's status to 'tallied' (the code even comments # Simulate Swarm Success for the demonstration), and evolve.py's apply_maxval_vector is an empty placeholder. A session audit identified this as the single largest "language ≠ code" gap in the repo — exactly the kind of "lies" the project's own CLAUDE.md and isomorphism doctrine reject.

What does this PR change?

A minimal, honest relabel — strings, docstrings, and comments only. Every misleading line now carries a [DEMO] marker and accurate wording (no real dispatch, no branch merge, simulated success). Behavior is unchanged: the loop still tallies features and terminates at [PERFECT]; the real HardwarePipingManager leasing/GC is untouched. No functions were renamed (callers cli.py / Makefile invoke these scripts unchanged). This makes the code truthful without faking capability; actually implementing a real swarm is left as a deliberate future choice.

Is this change appropriate for the core library?

No. Fork-specific scripts. Internal fork PR.

What alternatives did you consider?

  • Implement a real recursive LLM swarm — out of scope for an honesty fix: it's a major build requiring model/agent integration. Relabeling is the correct, safe step now; implementing is a separate future project.
  • Delete the demo loop — rejected: it's wired into cli.py/Makefile (make evolve) and is a working demonstration of the gate flow; honest labeling preserves its value.
  • Only fix the most egregious lines — initially did, then a reviewer flagged two remaining overclaims (Swarm mutation queued, achieved absolute MaxVal); folded those in so the relabel is complete.

Does this PR contain multiple unrelated changes?

No. One problem: the swarm scripts overclaim what they do. Two files, one purpose.

Existing PRs

  • Reviewed open + closed PRs. Related: PR feat: YOLO - The Chaos Monkey #12 ("Chaos Monkey") and the evolution-matrix PRs introduced this loop; none honestly labels it. No prior labeling PR exists.

Environment tested

Harness Harness version Model Model version/ID
Claude Code (CLI) not surfaced Claude Opus 4.8 orchestration
  • python -c "import ast; ast.parse(...)" → PARSE OK for both files.
  • Ran recursive_evolution.py against a temp feature_map.json (pending features): it still tallies each feature across epochs and terminates at [PERFECT], now emitting [DEMO] markers throughout — behavior preserved, output honest.
  • Verified the diff is strings/docstrings/comments only (no control flow, no function renames; signatures byte-identical base vs HEAD); grep confirms no DEMO-less overclaiming line remains.
  • Host: Windows 11 ARM64, Python 3.14.

New harness support

N/A.

Evaluation

N/A for skill evals. Functional: before, the loop printed real-work claims while flipping a flag; after, it labels itself a demonstration and prints honest [DEMO] lines, with identical control flow. Independently reviewed (diff-is-strings-only confirmed, no renames, parse + demo-run verified, remaining overclaims swept).

Rigor

  • Skills change — N/A.
  • Tested adversarially — a reviewer re-ran parse + a demo loop and grepped for residual overclaims.
  • Did not modify behavior-shaping content.

Human review

  • A human has reviewed the COMPLETE proposed diff before submission.

Summary by CodeRabbit

  • Chores
    • Updated internal development scripts to run in demonstration mode with simulated behavior instead of performing autonomous or mutating operations.
    • Runtime output and status logs now clearly indicate demo/simulated execution and show simulated progress and completion messages for testing and validation.

The evolve.py / recursive_evolution.py loop is marketed as an 'Infinite
Recursive LLM Swarm' but performs no LLM work: it deterministically flips
each feature's status to 'tallied'. Relabel output/docstrings/comments with
[DEMO] markers and accurate descriptions (no real dispatch, no branch merge,
simulated success). Behavior unchanged; no functions renamed.
- 'Swarm mutation queued' -> 'Simulated mutation queued' [DEMO]
- '[TERMINATE] achieved absolute MaxVal' -> honest demo-complete message
@coderabbitai

coderabbitai Bot commented Jun 6, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 5293cfed-c7c1-4e01-89e4-1f61f85176e1

📥 Commits

Reviewing files that changed from the base of the PR and between 85cc315 and 1033b6a.

📒 Files selected for processing (1)
  • scripts/recursive_evolution.py

📝 Walkthrough

Walkthrough

Two evolution scripts are converted to demonstration mode: module docstrings, runtime logs, and two function implementations now simulate baseline → apply → finalize flows without dispatching agents or mutating code; subprocess output from evolve.py is captured.

Changes

Evolution scripts to demonstration mode

Layer / File(s) Summary
evolve.py demonstration stub
scripts/evolve.py
Module docstring and function implementations clarify the script is a simulation placeholder. apply_maxval_vector() and finalize_effect() are replaced with no-op stubs that print demo messages instead of dispatching agents or mutating code.
recursive_evolution.py demonstration mode
scripts/recursive_evolution.py
Module docstring reframes the orchestrator as a deterministic demonstration harness. Subprocess invocation still runs the evolve.py stub (output captured), and runtime logs adopt [DEMO] prefixes for initialization, lease grant/deny, per-feature tallying, and termination messages.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Poem

🐰 In demo fields the scripts now play,
No swarms take flight, no code will sway,
They print their tales in gentle light,
Baseline, apply, finalize — all in sight,
A quiet show of simulated might.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 25.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: relabeling the swarm loop as a demonstration to fix misleading claims about LLM work.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/honest-swarm-labeling

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@scripts/recursive_evolution.py`:
- Line 39: The change in recursive_evolution.py adds capture_output=True to the
subprocess.run call that invokes evolve.py (subprocess.run([... , target_dir],
capture_output=True)), which suppresses evolve.py's stdout/stderr and violates
the "no behavioral changes" requirement; revert this by removing the
capture_output=True argument (or set it to False) so subprocess.run forwards
evolve.py's output to the parent process as before, ensuring the call in
recursive_evolution.py continues to run evolve.py with visible stdout/stderr.
- Line 39: The subprocess.run call invoking sys.executable with 'evolve.py' and
target_dir should explicitly pass check=False (i.e. subprocess.run([...],
capture_output=True, check=False)) to make it clear that non-zero exits are
intentionally ignored, or alternatively inspect the returned CompletedProcess
(assign result = subprocess.run(...)) and handle result.returncode/error output;
update the call in the subprocess.run invocation in recursive_evolution.py
accordingly.
- Line 45: The print statement uses an unnecessary f-string with no
placeholders; update the print call in scripts/recursive_evolution.py (the line
containing print(f"  [HARDWARE][DEMO] Lease denied (System fully utilized).
Simulated mutation queued for next cycle.")) to use a normal string literal
instead (remove the leading f so it becomes print("  [HARDWARE][DEMO] Lease
denied (System fully utilized). Simulated mutation queued for next cycle.")).
- Line 37: The print call in recursive_evolution.py uses an unnecessary
f-string: change the print invocation that currently reads print(f" 
[GATE][DEMO] Lease granted. Running the evolve.py stub (no real swarm)...") to
use a plain string literal instead (print("  [GATE][DEMO] Lease granted. Running
the evolve.py stub (no real swarm)...")), removing the leading f so there are no
unused f-string prefixes.
- Line 42: The print call prints a static string with no placeholders, so remove
the unnecessary f-string prefix from the print invocation (the line containing
print(f"  [HARDWARE] Releasing compute lease...")) and change it to a normal
string literal to avoid misleading readers and tiny runtime overhead.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: e6ea9770-4c5f-47ed-aae9-15384d4e7b1e

📥 Commits

Reviewing files that changed from the base of the PR and between 06b2637 and 85cc315.

📒 Files selected for processing (2)
  • scripts/evolve.py
  • scripts/recursive_evolution.py

Comment thread scripts/recursive_evolution.py Outdated
print(f" [GATE] Lease granted. Engaging swarm for hypothesis mutation...")
print(f" [GATE][DEMO] Lease granted. Running the evolve.py stub (no real swarm)...")
hw_manager.flush_vram() # VRAM time-slicing prep
subprocess.run([sys.executable, os.path.join(os.path.dirname(__file__), 'evolve.py'), target_dir], capture_output=True)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚖️ Poor tradeoff

Behavioral change contradicts PR objectives.

The PR objectives explicitly state "No behavioral or API changes" and "changes only strings, docstrings, and comments," but adding capture_output=True to subprocess.run() changes runtime behavior by suppressing the output from evolve.py that would otherwise be visible to the user. Previously, all output from the evolve.py script would flow directly to stdout/stderr; now it is captured and discarded.

🧰 Tools
🪛 Ruff (0.15.15)

[error] 39-39: subprocess call: check for execution of untrusted input

(S603)


[warning] 39-39: subprocess.run without explicit check argument

Add explicit check=False

(PLW1510)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/recursive_evolution.py` at line 39, The change in
recursive_evolution.py adds capture_output=True to the subprocess.run call that
invokes evolve.py (subprocess.run([... , target_dir], capture_output=True)),
which suppresses evolve.py's stdout/stderr and violates the "no behavioral
changes" requirement; revert this by removing the capture_output=True argument
(or set it to False) so subprocess.run forwards evolve.py's output to the parent
process as before, ensuring the call in recursive_evolution.py continues to run
evolve.py with visible stdout/stderr.

🧹 Nitpick | 🔵 Trivial | 💤 Low value

Add explicit check parameter to subprocess.run().

The call to subprocess.run() should include an explicit check=False parameter to clarify that non-zero exit codes are intentionally ignored, or add error handling if failures should be detected.

♻️ Proposed fix
-            subprocess.run([sys.executable, os.path.join(os.path.dirname(__file__), 'evolve.py'), target_dir], capture_output=True)
+            subprocess.run([sys.executable, os.path.join(os.path.dirname(__file__), 'evolve.py'), target_dir], capture_output=True, check=False)
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
subprocess.run([sys.executable, os.path.join(os.path.dirname(__file__), 'evolve.py'), target_dir], capture_output=True)
subprocess.run([sys.executable, os.path.join(os.path.dirname(__file__), 'evolve.py'), target_dir], capture_output=True, check=False)
🧰 Tools
🪛 Ruff (0.15.15)

[error] 39-39: subprocess call: check for execution of untrusted input

(S603)


[warning] 39-39: subprocess.run without explicit check argument

Add explicit check=False

(PLW1510)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/recursive_evolution.py` at line 39, The subprocess.run call invoking
sys.executable with 'evolve.py' and target_dir should explicitly pass
check=False (i.e. subprocess.run([...], capture_output=True, check=False)) to
make it clear that non-zero exits are intentionally ignored, or alternatively
inspect the returned CompletedProcess (assign result = subprocess.run(...)) and
handle result.returncode/error output; update the call in the subprocess.run
invocation in recursive_evolution.py accordingly.

Source: Linters/SAST tools

Comment thread scripts/recursive_evolution.py
Comment thread scripts/recursive_evolution.py Outdated
@ZuluYokohama

Copy link
Copy Markdown
Collaborator Author

@coderabbitai Thanks for the review. Addressed the F541 f-string nits on the two lines this PR changed ([GATE][DEMO] and [HARDWARE][DEMO]).

Respectfully declining two findings, with evidence:

  1. capture_output=True on the subprocess.run (flagged Major, "behavioral change"): this is pre-existing, not introduced by this PR. It is a context line in the diff (no +), and git show origin/master:scripts/recursive_evolution.py shows capture_output=True already present at line 36. This PR changes only strings/docstrings/comments — removing capture_output would itself be the behavioral change the PR avoids. Leaving it as-is.

  2. check=False suggestion: same pre-existing subprocess.run line, untouched by this PR — adding the kwarg would be scope creep onto a line this change doesn't modify. Deferring to a separate cleanup.

Behavior remains unchanged; the relabel is strings-only.

@coderabbitai

coderabbitai Bot commented Jun 6, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

@ZuluYokohama: Thanks for the detailed evidence — you're correct on both counts.

  1. capture_output=True: Confirmed — that line is pre-existing context, not introduced by this PR. Flagging it as a behavioral change introduced here was a mistake on my end. No action needed.

  2. check=False: Agreed that adding it to an untouched line would be scope creep. Deferring to a separate cleanup is the right call.

The F541 fixes are appreciated. The PR looks good — strings/docstrings/comments only with honest [DEMO] labeling throughout.

Would you like me to open a follow-up issue to track the check=False addition on that subprocess.run call, so it doesn't get lost?

(◕‿◕✿) 🐇

@ZuluYokohama
ZuluYokohama merged commit a6944a8 into master Jun 6, 2026
1 of 2 checks passed
@ZuluYokohama
ZuluYokohama deleted the fix/honest-swarm-labeling branch June 6, 2026 14:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants