New Scaling Axis for Generative Models: Beyond Parameters and Data
Harvard paper introduces a third scaling axis for generative models, achieving 4.1x FLOP, 6.2x sample, and 47% parameter efficiency, with 300x faster convergence on ImageNet. Cross-domain applicability (video, NLP, robotics) suggests potential to reshape pretraining strategies and reduce compute costs, creating startup opportunities in efficient training. This is a significant research breakthrough that could impact hardware demand and model development economics. New: Skaling paper fixes independence assumption in Chinchilla/Kaplan scaling laws, reducing MAPE by 1.5-3x and enabling accurate extrapolation with 10x less compute—complementary insight for efficient training and compute allocation.