0xBakeer/TandemLLM
Inference engine for Qwen3.8-27B on one DGX Spark: speculative draft trees sized by StairCut, NVFP4 kernels, exact recurrent-state caches
GitHub repository with 21 stars and 0 forks.
Language: Python
Topics: dgx-spark, llm-inference, nvfp4, qwen, speculative-decoding