Ternary compression and reproducible quantization push compact local models toward practical deployment
PrismML's Bonsai 2 is reported as a 5.95 GB ternary Qwen3.8-based model with Apache 2.0 licensing, Apple Silicon support, and a custom llama.cpp fork. Pinecone's MIT-licensed VQ-Bench strengthens the evaluation infrastructure for compression, but Bonsai's conflicting retention results and category-specific losses make independent, workload-specific testing essential.
Sources (2)
Updated Sep 18, 2026