Summary
vllm is the dominant open-source high-throughput inference/serving engine for LLMs. This repo has zero instrumentation for vLLM's offline, in-process Python execution API — the vllm.LLM class used for direct batch/local inference without an HTTP server. This is distinct from vLLM's separately-covered OpenAI-compatible server mode (which is out of scope here since it's just an HTTP endpoint that any OpenAI-client-based tracing already sees).
LLM.generate(prompts, sampling_params), LLM.chat(messages), and LLM.embed(prompts) are direct in-process function calls returning RequestOutput/pooling objects — there is no HTTP boundary for the existing OpenAI wrapper (py/src/braintrust/wrappers/openai.py) to intercept, so these calls are completely untraced today even for users who have braintrust.auto_instrument() enabled.
What needs to be instrumented
LLM.generate(prompts, sampling_params, ...) — batch text/chat completion execution, sync
LLM.chat(messages, sampling_params, ...) — chat-shaped completion execution
LLM.embed(prompts, ...) — embedding/pooling execution (for pooling models per vLLM's docs distinction between generative and pooling models)
- Async equivalents via
AsyncLLMEngine / AsyncLLM used by the vllm.entrypoints layer
- Streaming/enqueue-style APIs (
LLM.enqueue, LLM.wait_for_completion) where present in the pinned version
Weekly downloads
Weekly downloads: 526,872 (as of 2026-09-07; https://pypistats.org/packages/vllm)
Braintrust docs status
not_found — checked https://www.braintrust.dev/docs/integrations and https://www.braintrust.dev/docs/guides/tracing/integrations; neither lists vllm/vLLM. Braintrust's only documented vLLM-adjacent path treats vLLM purely as a self-hosted HTTP endpoint behind Braintrust's AI gateway/custom-provider config (per braintrust.dev/blog/any-framework-any-provider), which does not apply to the in-process LLM class execution path described above.
Upstream sources
Local repo files inspected
py/src/braintrust/integrations/ — no vllm/ directory (27 integration directories checked, none match)
py/src/braintrust/wrappers/ — no vLLM wrapper
py/pyproject.toml [tool.braintrust.matrix] — no vllm entry
py/noxfile.py — no test_vllm session
py/src/braintrust/integrations/versioning.py — no mention
- Repo-wide case-insensitive grep for
vllm under py/src/braintrust/ — no real matches (only incidental substring hits inside unrelated binary cassette fixture bytes, e.g. py/src/braintrust/integrations/litellm/cassettes/latest/test_litellm_image_generation.yaml:42)
Summary
vllmis the dominant open-source high-throughput inference/serving engine for LLMs. This repo has zero instrumentation for vLLM's offline, in-process Python execution API — thevllm.LLMclass used for direct batch/local inference without an HTTP server. This is distinct from vLLM's separately-covered OpenAI-compatible server mode (which is out of scope here since it's just an HTTP endpoint that any OpenAI-client-based tracing already sees).LLM.generate(prompts, sampling_params),LLM.chat(messages), andLLM.embed(prompts)are direct in-process function calls returningRequestOutput/pooling objects — there is no HTTP boundary for the existing OpenAI wrapper (py/src/braintrust/wrappers/openai.py) to intercept, so these calls are completely untraced today even for users who havebraintrust.auto_instrument()enabled.What needs to be instrumented
LLM.generate(prompts, sampling_params, ...)— batch text/chat completion execution, syncLLM.chat(messages, sampling_params, ...)— chat-shaped completion executionLLM.embed(prompts, ...)— embedding/pooling execution (for pooling models per vLLM's docs distinction between generative and pooling models)AsyncLLMEngine/AsyncLLMused by thevllm.entrypointslayerLLM.enqueue,LLM.wait_for_completion) where present in the pinned versionWeekly downloads
Weekly downloads: 526,872 (as of 2026-09-07; https://pypistats.org/packages/vllm)
Braintrust docs status
not_found— checked https://www.braintrust.dev/docs/integrations and https://www.braintrust.dev/docs/guides/tracing/integrations; neither listsvllm/vLLM. Braintrust's only documented vLLM-adjacent path treats vLLM purely as a self-hosted HTTP endpoint behind Braintrust's AI gateway/custom-provider config (per braintrust.dev/blog/any-framework-any-provider), which does not apply to the in-processLLMclass execution path described above.Upstream sources
LLMclass quickstart: https://docs.vllm.ai/en/stable/getting_started/quickstart/Local repo files inspected
py/src/braintrust/integrations/— novllm/directory (27 integration directories checked, none match)py/src/braintrust/wrappers/— no vLLM wrapperpy/pyproject.toml[tool.braintrust.matrix]— novllmentrypy/noxfile.py— notest_vllmsessionpy/src/braintrust/integrations/versioning.py— no mentionvllmunderpy/src/braintrust/— no real matches (only incidental substring hits inside unrelated binary cassette fixture bytes, e.g.py/src/braintrust/integrations/litellm/cassettes/latest/test_litellm_image_generation.yaml:42)