Frontier Architecture Innovations: Decoupled Knowledge and Reasoning
Intern-S2-Mobius decouples knowledge (FFN) and reasoning (self-attention), achieving 4x inference speedup and 37.4% less training data. This challenges monolithic transformers and could reshape scaling. Also, Pathway's 150M model breaks ARC-AGI-1 cost-efficiency Pareto frontier via latent recurrent reasoning. NVIDIA's Nemotron 3.5 Lightning (30B MoE, 3B active) adds an efficient open model for always-on agents. Matryoshka LM Suites train nested model families in one run, improving training efficiency.
Sources (3)
Updated Aug 21, 2026