Senior DevOps / MLOps Engineer | AI Infrastructure · GPU Platforms · Kubernetes · Linux Systems
I build infrastructure where GPU compute, high-performance networking, Kubernetes, storage, and ML workloads meet.
My current focus is AI Factory engineering — GPU platforms, cluster networking, Linux reliability, and production ML infrastructure.
- GPU Platforms — NVIDIA GPUs, CUDA, DeepStream, DCGM, training & inference
- AI Cluster Networking — InfiniBand, RDMA/RoCE, Clos, BGP/ECMP, NetBox
- Kubernetes — Kubernetes, CNI, GPU workloads, Helm, GitOps
- Linux Reliability — eBPF, cgroups, OOM analysis, kernel diagnostics
- MLOps — model lifecycle, evaluation, lineage, serving, observability
- Automation — AWS, Terraform, Ansible, GitHub Actions, CI/CD
Executable AI cluster network lab using NetBox, Containerlab, Clos fabrics, BGP/ECMP, failure injection, and automated validation.
NetBox · Containerlab · FRR · BGP · ECMP · Clos
Rust + eBPF diagnostic tool for finding hidden shared-memory and CUDA memory pressure behind Linux cgroup OOM failures.
Rust · eBPF · cgroups · CUDA · Linux
GPU video reasoning platform combining DeepStream perception, VLM verification, Kubernetes, and multi-stream inference.
DeepStream · CUDA · Kubernetes · VLM · FastAPI
AI infrastructure knowledge system covering GPU, RDMA/InfiniBand/RoCE, storage, distributed training, inference, MLOps, and performance engineering.
Go-based Linux kernel postmortem toolkit for validated vmcore analysis and evidence-backed crash diagnostics.
Go · Linux Kernel · vmcore · crash · drgn
Production-shaped computer vision lifecycle from dataset and training to evaluation, promotion, ONNX deployment, and drift monitoring.
PyTorch · CV · MLOps · ONNX · W&B
Active-learning pipeline turning production false positives into reviewed hard negatives, retraining data, and promotion signals.
YOLO · CLIP · Label Studio · W&B
GPU research harness comparing RAG, SFT, and RAFT across quality, latency, and VRAM trade-offs.
NVIDIA GPU · QLoRA · RAG · RAFT · LLM Evaluation
Cloudflare-native LLM knowledge graph and MCP retrieval backend for Obsidian and Markdown knowledge bases.
Cloudflare Workers · R2 · D1 · Vectorize · MCP
Reusable GitHub Actions for Kubernetes delivery and AI-assisted repository automation.
GitHub Actions · Kubernetes · GitHub Apps
Real-time AI English speaking coach for engineers using Gemini Live, structured feedback, authentication, and usage quotas.
- CapEx Lens — AI infrastructure CAPEX and hyperscaler economics dashboard
- Kernel Lens — A clearer view into Linux kernel development.
- AKBO — A lightweight, tablet-friendly sheet music viewer for piano player.
- NVIDIA AI Factory architectures
- GPU cluster operations
- InfiniBand / RDMA / RoCE
- NCCL and distributed GPU communication
- Kubernetes GPU orchestration
- NVIDIA GPU / Network Operators
- AI storage and data paths
- GPU / Linux performance engineering
Building AI infrastructure from the kernel to the cluster — and from the cluster to the model.





