professorpalmer/bonsai-ada-surgery
Bonsai 2 27B: 262k context (trained max) on 12 GB, 54 to 97 tok/s decode, 2x prefill, q8_0 KV recipe, grammar tool calls 9/9. Same ternary weights. Windows bundle + receipts.
GitHub repository with 11 stars and 0 forks.
Language: Python
Topics: bonsai, cuda, llama-cpp, local-llm, quantization, qwen, rtx-4070, speculative-decoding, ternary