Skip to content

No prompt caching on /provider/v1/chat/completions (GOAT): prompt_cache_key never sent, cache reads ~0% #132

Description

@maxitromer

Hi, first of all thanks for the extension — the GOAT transport work in 0.6.0 and the pricing fixes in 0.7.0 are much appreciated.

I'm opening this as a follow-up to #83. That issue nailed the pricing half (fixed in 0.7.0, confirmed) but the caching half was closed as "gateway ignores the key, nothing we can do client-side". I think there's new information that reopens it — specifically for the Provider API transport (/provider/v1/chat/completions), which is what GOAT accounts actually use. The #83 probe tested /alpha/generate, but that's the Go fallback, not the GOAT path.

Environment

  • pi-commandcode-provider 0.7.3 (npm)
  • pi 1.0.0 / pi-ai 1.0.0
  • Arch Linux, kernel 6.18.52-1-lts, node v24.13.0
  • CommandCode GOAT plan (so transport is always provider, never generate)
  • Model: meta/muse-spark-1.3-contributor (also affects the other Meta models and anything else on the openai-completions branch)

Symptom

Running long agentic sessions (lots of tool calls, many turns), the reported cache usage on commandcode models sits permanently under 1%. Same workload on another OpenAI-compatible provider in the same pi setup reports cache hits near 100%. So this isn't pi misreading the usage — the cache reads genuinely aren't happening, and every turn re-bills the full prefix at fresh-input price ($0.10/M instead of $0.002/M on contrib — roughly 50x). On long sessions that adds up fast.

Root cause (traced through the code)

For GOAT, transport.ts pins streamProvider — the request goes to POST https://api.commandcode.ai/provider/v1/chat/completions through pi-ai's native openai-completions streaming. And in that path, pi-ai's buildParams (pi-ai/dist/api/openai-completions.js) gates the cache key like this:

prompt_cache_key: (model.baseUrl.includes("api.openai.com") && cacheRetention !== "none") ||
    (cacheRetention === "long" && compat.supportsLongCacheRetention)
    ? clampOpenAIPromptCacheKey(options?.sessionId)
    : undefined,
prompt_cache_retention: cacheRetention === "long" && compat.supportsLongCacheRetention ? "24h" : undefined,

Three independent reasons this is always undefined for commandcode models:

  1. model.baseUrl is https://api.commandcode.ai/provider/v1 — never contains api.openai.com, so the first clause is dead.
  2. cacheRetention defaults to "short" (only "long" with PI_CACHE_RETENTION=long), so the second clause is dead by default.
  3. Even if a user sets PI_CACHE_RETENTION=long, the extension registers openai-branch models with supportsStore: false and no supportsLongCacheRetention (index.ts, createProviderConfig), so prompt_cache_retention stays undefined anyway.

On top of that, no other cache/stickiness signal is sent on this path: getCompatCacheControl requires cacheControlFormat === "anthropic" (never true on the openai branch), and session-affinity headers are only emitted for openrouter (sendSessionAffinityHeaders). So the request carries literally nothing the gateway could use for sticky routing to a warm Meta backend — and the Contrib pool without sticky routing is exactly what @Star-233 and @beyondhumanwork documented in #83 (coin-flip hits, fresh keys hitting other pairs' warm cache, i.e. no key participation in cache identity).

Why the #83 conclusion doesn't cover this path

The #83 A/B tested /alpha/generate and correctly found the gateway drops unknown fields there. But GOAT never touches /alpha/generate — transport.ts only falls back there on 403 upgrade_required. The Provider API is a different gateway surface, and per the last comment on #83 from @urawazakun, prompt_cache_key is forwarded on /provider/v1/responses (just not on /chat/completions): keyed requests on meta/muse-spark-1.3-contributor hit ~99% from the 2nd call vs 0/5 keyless. Small sample, but it points exactly where the fix should go.

Also worth noting: pi-ai's openai-responses branch sends prompt_cache_key unconditionally (cacheRetention === "none" ? undefined : clamp(sessionId) — no api.openai.com gate), which is consistent with that observation.

Suggested directions (whatever you think fits best)

  • a) Expose an openai-responses api mapping for the Meta models (and any other models CommandCode serves on /provider/v1/responses), so the request goes through the endpoint that actually forwards the key. apiForModelId currently only knows openai-completions vs anthropic-messages.
  • b) Alternatively, inject prompt_cache_key (stable per session, e.g. from options.sessionId) via onPayload on the existing chat/completions path, and verify whether CommandCode's gateway forwards it there too.
  • c) At minimum, declare supportsLongCacheRetention for models where CommandCode honors prompt_cache_retention, so PI_CACHE_RETENTION=long has an effect.

Happy to test any branch against the live GOAT endpoint and report back cache-read numbers, same as @beyondhumanwork did for #83. I can run A/B pairs (keyed vs control, immediate-repeat and isolated fixtures) and post the cacheReadTokens distributions.

Thanks again for maintaining this — and for the thorough #83 discussion that made this trace possible.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't workingproviderProvider-related issues

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions