Benchmark runs for keyline: every real-model run, kept whole (prompt, every turn and tool call, renders, summary), so a change is judged on runs, not guesses.
reference-ad/: Claude Code builds the reference ad through keyline; one folder per run,results.tsvacross runs.versus-browser/: the same model makes the same images with keyline, HTML + headless Chrome, and Playwright MCP. Its README is the benchmark page on keyline.dev.
The harness is keyline's tests (tests/llm_e2e.rs, tests/versus_browser/). Check this repo out beside keyline as keyline-bench, or point KEYLINE_BENCH at it, and runs land here. Run them from keyline's root; the commands are in each README.
Licensed like keyline: FSL-1.1-ALv2.