Open source C/C++ engine for high-performance local LLM inference and on-device AI.
-
Updated
Jul 14, 2026 - C++
Open source C/C++ engine for high-performance local LLM inference and on-device AI.
Data-only launcher contracts and runtime-readiness surfaces for PCCX™ model integration.
Portable, provider-neutral model invocation contracts, adapters, and conformance tools
A self-contained native runtime and model format for local AI video generation.
GPT-OSS B20 Local Execution. Lightweight local environment for running it with Python 3.12 and CUDA acceleration. - Run GPT-OSS B20 entirely offline - Optimize text generation with GPU - Enable fast, secure inference on consumer hardware.
A complete inference runtime for open-weight large language models, enabling efficient execution through streaming weights, quantization, and memory-aware scheduling
Yollama | "Your-Ollama! Take control with true data privacy, plus, real offline Freedom for you and your LLM AI model inferencing!
To associate your repository with the model-runtime topic, visit your repo's landing page and select "manage topics."