feat(models): use the Provider API OpenAI Responses wire - #133
pierreraby wants to merge 2 commits into
Conversation
Command Code documents /provider/v1/responses for OpenAI and open models, and /provider/v1/models advertises the routes that serve each model through its supported_endpoints field. Resolve the wire per model from that field: Claude stays on /v1/messages, models advertising /responses use the OpenAI Responses wire, and every other model keeps /v1/chat/completions. The catalog is the source of truth, so a model whose entry omits the field falls back to Chat Completions, which every non-Claude model serves. The model cache now stores the resolved wire and bumps to version 2.
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.Scanned FilesNone |
Memory benchmarkCompared base
Interpretation: No metric shows a clear Base-to-PR difference beyond its measured run-to-run variation. Extension overhead above pi baseline
Values are medians of 6 alternating, paired runs. The value after Environment: pi
|
|
@patlux @karaaslanz — quick nudge on this one. CI is green, and the fix is now field-confirmed: same contributor session went from $0.0112 (108k input, 0% cache on Chat Completions) to $0.00084 (113k input, ~95-99% cached on Responses) — ~13x cheaper at equal workload (details on #132). Every long session without this bills the full prefix, so a timely review would save real money. Thanks! |
Carried forward from #124 — all credit to @Mario-pereyra for the design and the first commit, which keeps his authorship here.
Why a new PR rather than updating #124: its base predates #127, and the stale Grok pricing trips the date-guard as of today, so its CI can no longer go green untouched; pushing to a contributor fork didn't feel appropriate either. This branch is #124 rebased on current main, plus one mock-test fix (below).
Verified since #124 (live, on a GOAT plan):
function_callemitted, tool result fed back, assistant answer incorporates it./responseswithprompt_cache_key: 1251/1333 input tokens reported cached on immediate repeat (~94%). Same setup on/chat/completions: 0 cached even with an explicit key. This is the No prompt caching on /provider/v1/chat/completions (GOAT): prompt_cache_key never sent, cache reads ~0% #132 fix — repeated prefixes will bill at cache-read instead of full price.test-pi-local.mjsPASS including the Responses routing test and the RPC lifecycle test.The mock-test fix: the extra catalog model broke the RPC lifecycle count assertions (3, then 4), so it is now served only during its own test, with the shared models cache reset beforehand.
Cache misses bill the full prefix every turn, so this one was worth fast-tracking — happy to adjust anything.
Closes #132. Supersedes #124.