LoopFormer studies recurrent computation and adaptive inference depth in Qwen/Qwen2.5-0.5B-Instruct.
Can a recurrent language model learn useful iterative computation, and can inference determine when additional recurrent computation is useful, unnecessary, or harmful?
Reusing a learned block lets inference vary computational depth without adding a separate set of weights for every step. The project first tests whether recurrence executes an interpretable algorithm, then studies depth generalization, overscaling dynamics, and eventually compute-quality tradeoffs through stopping. Pointer chasing is the controlled starting environment: fresh rules live in each prompt, and exact intermediate targets make errors and progress checkable.
prompt
↓
frozen prelude
↓
shared recurrent block × T
↓
frozen coda
↓
readout
One recurrent block and its LoRA adapters are reused across all loops. Original pretrained weights stay frozen. The prelude runs once; the frozen coda reads each loop's state for supervision and evaluation. The next loop consumes the recurrent hidden state, rather than the coda output or a decoded answer. See architecture.
| Evidence | Scope and limits |
|---|---|
| One-loop equivalence, shared-weight and gradient-scope tests | Validated model surgery; useful task execution needs separate evidence. |
| Fresh 30k depth-6 experiment | Strong unseen-mapping trajectories and bounded extension beyond training depth; one training run on development evaluation sets. |
| Joint learned-completion ablation | Stop timing and pointer execution do not extend as well as the CE-only reference; one development run. |
| Evaluation and reproducibility | Seeded data, exact intermediate supervision, full per-loop exports, source/data/checkpoint provenance, selection discipline, synchronized timing and throughput. |
| Recorded contract-test validation | Architecture, data, training and metric contracts; tests do not establish empirical research gates. |
For the fresh 30k run, step 2500 achieved the following complete-trajectory accuracy (every nominal intermediate state correct):
| Task depth | 1–6 | 7 | 8 | 9 | 10 | 11 | 12 | 13–16 |
|---|---|---|---|---|---|---|---|---|
| Accuracy | 98.0% | 94.4% | 94.4% | 73.6% | 41.6% | 14.4% | 1.6% | 0% |
Depths 1–6 aggregate 750 examples; each individual depth has 125. Step 2500 is the working depth-generalization checkpoint among three fully evaluated candidates; step 3250 remains the trained-loss-selected reference. Seed-17 evaluations are development diagnostics, and seed 29 remains reserved for confirmation. This supports execution beyond the training depth followed by degradation at greater task depths. It does not establish arbitrary-depth execution, general reasoning, or post-completion overscaling damage.
The completed mechanism diagnostics and learned-completion ablation narrow the failure but do not identify its internal cause. Current status records the evidence and remaining gates.
Training supports W&B logging to the loopformer project, with local metrics retained. Standalone evaluations also support W&B, including publishing saved results without rerunning inference.
The checkpoint comparison is complete; later training reduced depth generalization. The next diagnostic is bash probe_depth12.sh on the CUDA desktop. See results and run instructions.
The loop-balanced loss ablation and joint completion training did not improve farther-depth generalization. Further experiments should target a specific unresolved mechanism. Terminal overscaling evaluation already provides repair/damage and continuous-survival metrics, but pretrained terminal sweeps remain deferred. Asymmetric dynamics and transfer remain planned.
The next major direction is adaptive inference compute. Offline oracle/heuristic analysis, opt-in stopped inference, synchronized latency recording, and a gated lightweight head trainer now have code and tests. The adaptive-compute guide is the single source for methods and usage; no pretrained adaptive policy or speedup result has been verified. The original continuing-pointer tasks need explicit completion semantics before answer changes can be interpreted as damage.
Start with the research plan, current status, evaluation guide, and documentation index.
Use Python 3.11 and run commands from the repository root. Windows/NVIDIA users should follow the WSL2/CUDA setup guide first; other environment details are in setup.
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
python -m pip checkDownload the base revision used for the reproducible experiments:
hf download Qwen/Qwen2.5-0.5B-Instruct --revision 7ae557604adf67be50417f59c2c2f167def9a775Check ordinary model loading and generation:
python scripts/smoke_test_qwen.pyActivate .venv in each new terminal. The commands below default to CPU unless a device is supplied. Use --device mps on Apple Silicon or --device cuda on the configured NVIDIA desktop. For deterministic CUDA runs, set:
export CUBLAS_WORKSPACE_CONFIG=:4096:8Dataset code lives in scripts/dataset/. Preview five examples without writing files:
python -m scripts.dataset --seed 17 --dry-runUse --dry-run 10 for ten examples. Reproduce the existing dataset with master seed 17, the pinned tokenizer above, and this complete configuration:
python -m scripts.dataset \
--seed 17 \
--train-count 10000 \
--validation-count 1000 \
--test-count 1000 \
--depth-test-count 1000 \
--min-depth 1 \
--max-train-depth 8 \
--max-eval-depth 16 \
--output data/pointer/seed-17
python -m scripts.dataset --verify data/pointer/seed-17The existing seed-17 dataset already reaches depth 16 for development evaluation. The current fresh-training experiment uses a separate 30k seed-37 dataset. Generated files are excluded from Git, so use the reproduction command above only if the dataset is missing on a new machine.
File under data/pointer/seed-17/ |
Examples | Depths | Examples per depth |
|---|---|---|---|
train.jsonl |
10,000 | 1–8 | 1,250 |
validation.jsonl |
1,000 | 1–8 | 125 |
test.jsonl |
1,000 | 1–8 | 125 |
depth_test.jsonl |
1,000 | 9–16 | 125 |
Current training uses the separate seed-37 30k dataset at depths 1–6. validation_max_depth: 8 controls monitoring; the full evaluator also reads seed-17 depth_test.jsonl at depths 9–16. See dataset and training details.
Inspect three ordinary-model questions before running the full test:
python -m scripts.eval.naive_test --model Qwen/Qwen2.5-0.5B-Instruct --test
python -m scripts.eval.naive_test --model Qwen/Qwen2.5-0.5B-InstructThe ordinary baseline uses three-shot examples at depths 1, 2, and 3, defined in prompts/pointer_task.txt. --model also accepts a saved model directory. Models load locally by default; --download permits missing downloads.
Progress and throughput appear in the terminal. Predictions and summaries go to eval/pointer_task/. See baseline evaluation for loading and scoring options.
On the CUDA desktop, prepare the new depth-independent dataset and check the isolated executor pipeline:
bash prepare_executor.sh --dry-run
bash prepare_executor.sh
bash train_executor.sh --dry-run
bash train_executor.sh --smoke-testAfter inspecting the smoke results, run bash train_executor.sh, then
bash eval_executor.sh. The default trains all recurrent-block weights with an
isolated controller and per-loop supervision. Full-model and LoRA controls are
also available. See the current run instructions
for configuration, exact data reproduction and output locations.
A compact Rich dashboard shows progress, ETA, losses and memory. W&B uses project
loopformer. Checkpoints and logs go to models/stage1_pointer/; model binaries
are excluded from Git. CUDA fit, speed and quality need the new desktop smoke/run.
bash eval_executor.sh runs matched-count/precision diagnostics, full-loop evaluation and actual learned stopping on development data. Results and terminal logs go under eval/pointer_diagnostics/. The older depth-12 scripts remain available for historical runs.
Pass a complete saved step directory, including its adapter weights and tokenizer:
python -m scripts.eval.loop_test \
--model models/stage1_pointer/executor_r-seed61/step-000750 \
--data data/pointer/seed-61-independent/validation.jsonl \
--device cuda --loops 12 --testThis requires the complete checkpoint on the training desktop; local metadata alone is insufficient. Remove --test for all 1,536 validation queries. Outputs go to eval/pointer_loops/. This reads the model after every recurrent loop; the ordinary three-shot prompt is not used. See full-loop evaluation for commands and metrics.
Preview the separate absorbing-terminal task variant without loading a model or writing files:
python -m scripts.eval.overscaling_test --dry-runCheckpoint sweeps, repair/damage scoring, and survival exports are available but running them is deferred until further Stage 1 training and review. See overscaling usage.