Bizuayeu/GLM-5.3-Flash-NVFP4-2x-DGX-Sparks-BIZ
NVFP4 BIZ: run nvidia/GLM-5.3-Flash-NVFP4 on two DGX Spark-class GB10 nodes: vLLM TP=2, Marlin W4A16, FA2 prefill, MTP, prefix caching, 256K + images, bit-reproducible output; up to 45 tok/s decode with the NVFP4 BIZ AXL weights. Apache-2.0 code, MIT weights fetched separately. BIZ = business-use intent, not support or certification.
GitHub repository with 6 stars and 1 forks.
Language: Python