-
Nokia
- Munich, Germany
- https://anubhabbanerjee.github.io/
- https://linkedin.com/anubhab-banerjee/
Popular repositories Loading
-
Annotated-LLM-Runtime
Annotated-LLM-Runtime PublicFrom-scratch, heavily-annotated CUDA inference runtime for Qwen2.5-Coder-7B on H100 (sm_90). Custom INT4 packer, fused GEMV, paged KV, split-KV attention, CUDA graph decode — every hot path comment…
-
WarpGroup-backend
WarpGroup-backend PublicA high-performance C++ backend for extreme-context LLM inference. It replaces item-count batching with dynamic, VRAM-aware First-Fit Decreasing (FFD) bin packing. By using PyBind11 for async queuei…
-
VRAM-Conductor
VRAM-Conductor PublicC++17 + CUDA orchestrator for running multiple LLM agents on one constrained GPU. `lmxd` daemon does NVML-seeded admission control to stop llama.cpp OOM crashes; `LayerStreamer` + `PinnedHostPool` …
C++ 3
-
inter-llm-tokf
inter-llm-tokf PublicInter-LLM knowledge handover framework leveraging Open Knowledge Format (OKF) for Qwen family. Bypasses text paring and tokenization by passing binary token arrays directly into model embedding lay…
If the problem persists, check the GitHub status page or contact support.
