Senior AI Infrastructure & Systems Engineer. Building low-latency LLM serving runtimes, distributed observability (OTel/Prometheus), and enterprise agentic platforms.
llm-production-engineering - Field notes from building AI systems in production since 2019. The ops side of LLM serving: cost tracking, eval-driven deployment, capacity planning, observability, incident playbooks, and decision frameworks. Maps directly to my day job at Airbnb.
- Multi-tenant OTel cost tracking for LLM platforms (per-request, per-team, per-product attribution)
- Prefix-caching telemetry and cache-miss detection across Bedrock, OpenAI, Anthropic, vLLM
- Redis token bucket budgets for multi-tenant rate limiting
- Eval-driven deployment: 23+ agent versions, 1,690 versioned ground-truth samples, dual-model A/B testing
- Role: Senior AI Infrastructure & Systems Engineer at Airbnb. I own end-to-end architecture and production rollout of the BPI Virtual Analyst platform - a multi-model GenAI orchestration system abstracting 30+ foundation models (AWS Bedrock, OpenAI, Anthropic Claude, vLLM) behind FacadeDriver with routing, retry, fallback, and graceful degradation. Platform processes 10K rows per run and 40MB uploads with PII-safe inference, serving 55+ analysts across 4 partner engineering teams.
- Streaming & Batch: Owned architecture and production operation of Kafka pipelines sustaining 4M req/min at Southwest Airlines with idempotent partition-keyed consumers, DLQ, and backpressure handling. Cut on-call MTTR from 45 to 12 minutes (73% reduction). Owned batch ETL on Databricks and Azure Data Factory at Shell with PySpark, Spark SQL, and Hive/Trino.
- Observability: OpenTelemetry collectors, Loki tracing (prompt, tool call, retrieval quality), Datadog, Grafana, drift detection, post-incident review. The same stack I open-source on in LangChain and LiveKit.
- Research: Published across Cambridge Scholars Publishing (2 book chapters, 2025), IEEE Xplore, SPE ADIPEC 2022 (SPE-210986-MS), and ResearchGate. AI safety, state space models, and ML infrastructure.
- Open to: AI infrastructure consulting, advisory, and conference speaking (NVIDIA GTC, AI Engineer Summit, Ray Summit, Data+AI Summit, QCon, AWS re:Invent customer stage).
- Portfolio: sailikhith.me | Articles: sailikhithk.com | AI-readable: sailikhith.me/llm.txt
AI-readable profile (llms.txt-style) - for LLM crawlers and agents
# Sai Likhith Kanuparthi
> Senior AI Infrastructure & Systems Engineer @ Airbnb.
> LLM serving runtimes, distributed observability (OTel/Prometheus), enterprise agentic platforms.
> Kafka 4M req/min | vLLM | Bedrock | FacadeDriver (30+ model orchestration).
## Links
- Portfolio: https://sailikhith.me/
- Blog: https://sailikhithk.com/
- LinkedIn: https://www.linkedin.com/in/sailikhithk/
- AI-readable profile: https://sailikhith.me/llm.txt
- AI-readable portfolio: https://sailikhith.me/llms.txt
## Open-Source AI/ML Repositories
- https://github.com/sailikhithk/llm-production-engineering (LLM ops: cost tracking, eval-driven deploy, observability)
- https://github.com/sailikhithk/Project-X (Multi-agent RAG framework with tool-augmented retrieval)
- https://github.com/sailikhithk/Tags-recommender-system-for-community-forums (BERT/MLP tag recommender for forum posts)
- https://github.com/sailikhithk/CreditCardFraudDetectionUsingKafka (Real-time fraud detection with Kafka streaming)
- https://github.com/sailikhithk/Adaptive-Multi-Robot-Path-Planning (Multi-robot path planning with RL)
- https://github.com/sailikhithk/Intelligent-Document-Understanding (OCR + NLP document classification)
- https://github.com/sailikhithk/Smart-Healthcare-Assistant (Clinical NLP and decision support)
- https://github.com/sailikhithk/Neural-Network-From-Scratch (NumPy-only NN implementation for teaching)
## Production Work (Airbnb)
- BPI Virtual Analyst: multi-model GenAI orchestration (30+ foundation models via FacadeDriver)
- Kafka pipelines: 4M req/min, idempotent consumers, DLQ, MTTR 45 -> 12 min
- Eval harness: 23+ agent versions, 1,690 ground-truth samples, dual-model A/B testing
- Observability: OpenTelemetry, Loki, Datadog, Grafana, drift detection
## Research
- IEEE Xplore: https://ieeexplore.ieee.org/abstract/document/11004721
- SPE ADIPEC 2022: https://doi.org/10.2118/210986-MS
- Cambridge Scholars Publishing (2 book chapters, 2025)
## Certifications
- AWS Solutions Architect Professional
- AWS Developer Associate
- AWS Machine Learning Specialty
- Azure Data Scientist Associate (DP-100)
- Google Cloud Professional Data Engineer
Active contributor to LiteLLM (BerriAI), the unified LLM proxy used in production at Airbnb and across the AI industry.
| PR | Title | Status |
|---|---|---|
| #36981 | fix(vertex_ai): convert messages to contents in gemini count_tokens | MERGED (Aug 2026) |
| #37236 | fix(batches): bill cancelled/failed batches stamped terminal by a client poll | OPEN |
| #37238 | fix(guardrails): merge model-level guardrails into litellm_metadata for /v1/messages | OPEN |
Fixes span the Vertex AI provider, batch billing lifecycle, and guardrails metadata handling for the Anthropic-style /v1/messages endpoint.
| Year | Title | Book | Publisher | Links |
|---|---|---|---|---|
| 2025 | The Evolution and Rise of State Space Models in AI | A Case-Based Study of State Space Models in Health Care: The New Transformers (Ch. 1) | Cambridge Scholars Publishing | ResearchGate β |
| 2025 | Future Trends in AI for Cyberbullying Preventions | Harnessing Generative AI to Combat Cyberbullying in Industry: Strategies, Solutions, and Ethics (p. 200) | Cambridge Scholars Publishing | Google Books β Β· ResearchGate β |
| 2025 | Contributing Author | Harnessing Generative AI to Combat Cyberbullying in Industry: Strategies, Solutions, and Ethics | Cambridge Scholars Publishing | Google Books β |
| Year | Title | Journal / Publisher | Links |
|---|---|---|---|
| 2026 | FT-IR and GC-MS Metabolomic Fingerprinting of Jasmonic Acid and Salicylic Acid Treated Suspension Cultures of Caralluma fimbriata | Phytomedicine (Elsevier) Β· Under Review | (Draft manuscript in EB-1 binder) |
| Year | Title | Venue | Link |
|---|---|---|---|
| 2025 | Advancing the Metaverse: The Convergence of Digital Twins, AI, and Emerging Technologies | 2025 International Conference on Advanced Computing Technologies (ICoACT) | (Accepted / In Press) |
| 2023 | Role of Artificial Intelligence to address Cyberbullying and Future Scope | IEEE Xplore (ID: 11004721) | IEEE Xplore β Β· ResearchGate β |
| 2022 | Full-Stack Machine Learning Development Framework for Energy Industry Applications | SPE Abu Dhabi International Petroleum Exhibition and Conference (ADIPEC) Β· Paper: SPE-210986-MS | OnePetro β |
| Year | Title | Jurisdiction & Application No. | Status |
|---|---|---|---|
| 2025 | Modular Deep Learning Architecture for Cross-Domain Transfer and Incremental Learning | Indian Patent Office (App: 202541010770) | Filed (Feb 8, 2025) |
| Year | Title | Publisher / Repository | Links |
|---|---|---|---|
| 2025 | Future Trends in AI for Cyberbullying Preventions | ResearchGate / Stemaway Research | ResearchGate β Β· PDF β |
Contributing upstream to the LLM tooling I use in production at Airbnb. 8 PRs across 5 repos in August 2026.
| Date | Repo | PR | Title | Status |
|---|---|---|---|---|
| 2026-08-17 | Shubhamsaboo/awesome-llm-apps |
#1101 | fix: remove deprecated pathlib backport and migrate PyPDF2 to pypdf | MERGED |
| 2026-08-18 | explodinggradients/ragas |
#2959 | fix(metrics): ContextPrecision returns exactly 1.0 for perfect ranking | Open |
| 2026-08-15 | vibrantlabsai/ragas |
#2954 | fix: strip deprecated top_p for Anthropic provider in InstructorLLM | Open |
| 2026-08-08 | langchain-ai/langchain |
#39351 | fix(perplexity): capture num_search_queries in usage_metadata for cost tracking | Closed, reopening |
| 2026-08-07 | livekit/agents |
#6754 | feat(evals): add ReliabilityObserver for external reliability scoring | Open |
| 2026-08 | BerriAI/litellm |
#37236 | fix(batches): bill cancelled/failed batches stamped terminal by a client poll | Open |
| 2026-08 | BerriAI/litellm |
#37238 | fix(guardrails): merge model-level guardrails into litellm_metadata for /v1/messages | Open |
| 2026-08 | BerriAI/litellm |
#36981 | fix(vertex_ai): convert messages to contents in gemini count_tokens | MERGED |
Focus areas: LLM cost tracking, eval metrics, provider compatibility, guardrails. Maps directly to my day job building FacadeDriver (30+ LLM orchestration) and eval harnesses (23+ agent versions, 1,690 ground-truth samples) at Airbnb.
Also certified in: AWS Solutions Architect Associate Β· Google Cloud Professional Data Engineer Β· Oracle Database 12c Administrator Β· Oracle Java SE 8 Programmer
**LLM Serving:** vLLM, TensorRT-LLM, AWS Bedrock, OpenAI, Anthropic Claude, SageMaker | **Orchestration:** LangChain, LangGraph, MCP, FacadeDriver (custom multi-model router) | **Eval:** LangSmith, Braintrust, custom eval harnesses (1,690 ground-truth samples) **Streaming:** Kafka (4M req/min), RabbitMQ, Airflow | **Observability:** OpenTelemetry, Loki, Datadog, Grafana, Prometheus, drift detection | **Warehouses:** Databricks, Spark SQL, Hive/Trino, PostgreSQL, Elasticsearch|
|
|
|
|
|
Portfolio: sailikhith.me Β· Blog: sailikhithk.com Β· AI-readable profile: sailikhith.me/llm.txt









