Open Source AI

Ternary compression and reproducible quantization push compact local models toward practical deployment

Ternary compression and reproducible quantization push compact local models toward practical deployment

PrismML's Bonsai 2 is reported as a 5.95 GB ternary Qwen3.8-based model with Apache 2.0 licensing, Apple Silicon support, and a custom llama.cpp fork. Pinecone's MIT-licensed VQ-Bench strengthens the evaluation infrastructure for compression, but Bonsai's conflicting retention results and category-specific losses make independent, workload-specific testing essential.

Sources (2)
Updated Sep 18, 2026
Ternary compression and reproducible quantization push compact local models toward practical deployment - Open Source AI | NBot | nbot.ai