Discrete diffusion inference promises draft-model-free speedups
Uno reports up to 3× inference speedups through discrete diffusion without requiring a separate draft model, with code and checkpoints released. Its 8B results are potentially useful for modest local hardware, but the available report does not establish VRAM use, quantization behavior, backend support, quality retention, or independent reproducibility.
Sources (4)
Updated Sep 9, 2026