Turbo Ultimate Field Fare is a MacOS app that lets users run models like Qwen, Gemma, and GPT-OSS models with expert-streaming, allowing for large models on devices without a lot of memory.
-
Updated
Sep 11, 2026 - Swift
Turbo Ultimate Field Fare is a MacOS app that lets users run models like Qwen, Gemma, and GPT-OSS models with expert-streaming, allowing for large models on devices without a lot of memory.
Quality-first local MiniMax H3 video-series generation and loopback API for dual RTX 4090 workstations, with native audio, references, P8/P9 continuity, and preserved artifacts.
Research preview for reproducible MoE routing and memory oversubscription on consumer GPUs. Not a production inference engine.
Flash-backed mixture-of-experts inference on Apple devices with Swift/MLX.
DualDeadline adds separate gate/up and down-projection transfer deadlines to exact MoE offloading, with H200 validation and Triton-optimized predictors.
Stream PyTorch model blocks from NVMe or pinned CPU memory with a bounded GPU residency budget. LoRA-finetune models far larger than host RAM on one GPU - Qwen3-VL-32B and gpt-oss-120b train in under 6 GB of VRAM.
Fits BF16 models too large for your GPU's VRAM by streaming losslessly compressed weights layer by layer, with speculative decoding, and an experimental benchmark suite showing it beats AirLLM and Hugging Face Accelerate at matched memory.
Research artifact for Memory-Sovereign Inference: Output-Exact Execution Beyond Full Residency
To associate your repository with the model-offloading topic, visit your repo's landing page and select "manage topics."