Skip to content

feat(models): use the Provider API OpenAI Responses wire - #133

Open
pierreraby wants to merge 2 commits into
mainfrom
feat/openai-responses-carry
Open

pierreraby wants to merge 2 commits into
mainfrom
feat/openai-responses-carry

Conversation

@pierreraby

Copy link
Copy Markdown
Collaborator

Carried forward from #124 — all credit to @Mario-pereyra for the design and the first commit, which keeps his authorship here.

Why a new PR rather than updating #124: its base predates #127, and the stale Grok pricing trips the date-guard as of today, so its CI can no longer go green untouched; pushing to a contributor fork didn't feel appropriate either. This branch is #124 rebased on current main, plus one mock-test fix (below).

Verified since #124 (live, on a GOAT plan):

  • Tool-call round-trip on the Responses wire with deepseek/deepseek-v4.1-flash: function_call emitted, tool result fed back, assistant answer incorporates it.
  • Prompt caching on /responses with prompt_cache_key: 1251/1333 input tokens reported cached on immediate repeat (~94%). Same setup on /chat/completions: 0 cached even with an explicit key. This is the No prompt caching on /provider/v1/chat/completions (GOAT): prompt_cache_key never sent, cache reads ~0% #132 fix — repeated prefixes will bill at cache-read instead of full price.
  • Full unit suite, typecheck, prettier green; test-pi-local.mjs PASS including the Responses routing test and the RPC lifecycle test.

The mock-test fix: the extra catalog model broke the RPC lifecycle count assertions (3, then 4), so it is now served only during its own test, with the shared models cache reset beforehand.

Cache misses bill the full prefix every turn, so this one was worth fast-tracking — happy to adjust anything.

Closes #132. Supersedes #124.

Mario-pereyra and others added 2 commits October 4, 2026 01:05
Command Code documents /provider/v1/responses for OpenAI and open models,
and /provider/v1/models advertises the routes that serve each model through
its supported_endpoints field. Resolve the wire per model from that field:
Claude stays on /v1/messages, models advertising /responses use the OpenAI
Responses wire, and every other model keeps /v1/chat/completions.

The catalog is the source of truth, so a model whose entry omits the field
falls back to Chat Completions, which every non-Claude model serves. The
model cache now stores the resolved wire and bumps to version 2.
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Memory benchmark

Compared base 5783e9f with PR head 226e0dd on the same GitHub-hosted darwin runner. Lower values are better.

Metric Base PR PR − Base Change
Stable RSS 120.4 MiB ± 0.8 MiB 120.4 MiB ± 1.0 MiB -0.3 MiB ± 1.3 MiB -0.3%
Physical footprint 66.2 MiB ± 0.5 MiB 66.0 MiB ± 0.7 MiB -0.5 MiB ± 1.1 MiB -0.8%
Physical peak 76.4 MiB ± 1.1 MiB 76.2 MiB ± 1.0 MiB +0.1 MiB ± 0.7 MiB +0.2%

Interpretation: No metric shows a clear Base-to-PR difference beyond its measured run-to-run variation.

Extension overhead above pi baseline

Metric Base overhead PR overhead Difference
Stable RSS +6.1 MiB ± 1.3 MiB +5.9 MiB ± 1.1 MiB -0.3 MiB ± 1.3 MiB
Physical footprint +4.2 MiB ± 0.4 MiB +3.1 MiB ± 0.3 MiB -0.5 MiB ± 1.1 MiB
Physical peak +4.2 MiB ± 1.0 MiB +4.5 MiB ± 0.5 MiB +0.1 MiB ± 0.7 MiB

Values are medians of 6 alternating, paired runs. The value after ± is the median absolute deviation (MAD). Each process was sampled 12 times after a 2500 ms warm-up.

Environment: pi 0.84.4, Bun 1.4.0, darwin arm64. RSS comes from ps; physical footprint and peak come from macOS footprint.

This is a comparative signal, not a pass/fail threshold. GitHub-hosted runner noise can affect absolute values.

@pierreraby

Copy link
Copy Markdown
Collaborator Author

@patlux @karaaslanz — quick nudge on this one. CI is green, and the fix is now field-confirmed: same contributor session went from $0.0112 (108k input, 0% cache on Chat Completions) to $0.00084 (113k input, ~95-99% cached on Responses) — ~13x cheaper at equal workload (details on #132). Every long session without this bills the full prefix, so a timely review would save real money. Thanks!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

No prompt caching on /provider/v1/chat/completions (GOAT): prompt_cache_key never sent, cache reads ~0%

2 participants