AI Breakthrough Digest

Inference-Time Recurrence and Selective Computation Improve AI Efficiency

Inference-Time Recurrence and Selective Computation Improve AI Efficiency

Google DeepMind's Recirculation reuses hidden states at inference time without additional training, reporting a 23% perplexity reduction and 21% GSM8K improvement on Gemma 3. ShallowStream reports 52.1x per-frame prefill and 11.9x end-to-end latency reductions for streaming video by using shallow-layer KV caches, reinforcing the broader shift toward selective inference computation; quality, transfer, and deployment robustness remain open.

Sources (3)
Updated Sep 8, 2026