Skip to content
#

flash-decoding

Here are 3 public repositories matching this topic...

C++/CUDA inference engine for Qwen2.5-Coder-0.5B, written from scratch: INT4 GEMV/GEMM kernels, Flash-Decoding, a 24-layer decode engine and lossless speculative decoding. Verified against Hugging Face; benchmarked against bandwidth ceilings, PyTorch and llama.cpp.

  • Updated Sep 28, 2026
  • Python

Add this topic to your repo

To associate your repository with the flash-decoding topic, visit your repo's landing page and select "manage topics."

Learn more