This repo tracks benchmarks for Command Code, comparing it against other AI coding harnesses (like Claude Code) on cost, token usage, timing, and output quality, so we can see how much the harness itself (not just the underlying model) affects real-world results.
- Kimi K3: Command Code vs. Claude Code Harness: same model (Kimi K3) run through two different harnesses, comparing turns, tokens, timing, and cost.
- slash-design-showcase: benchmark and showcase for Command Code's
/designfeature.