4-Bit Optimal, Sub-1-Bit Falls Short
- 4-bit precision is almost universally optimal for zero-shot accuracy at fixed model bits, based on 35,000+ experiments across model families.
-...
Created by Collin Laird
Cutting edge research on model architectures, training methods, optimization, scaling and benchmarks
Explore the latest content tracked by ML Breakthroughs Digest
G-MIXER introduces geodesic mixup-based implicit semantics as a training-free approach to zero-shot compositional image retrieval, contrasting recent methods that rely on MLLMs for generating target descriptions.
A new review examines CNNs, LSTMs, GRUs, ConvLSTMs, Transformers, and Graph Neural Networks, assessing their distinct roles in processing spatiotemporal data for ground deformation tasks.
Three papers released the same day reveal a clear shift from text prompts to direct spatial interfaces.
Thinking Mode Fusion (TMF) unifies concise and long-form reasoning in one model for mathematical problem solving.
SMELT loops middle MoE layers twice while matching FLOPs, parameters, and KV cache, yielding 6.8–18% training efficiency gains and stronger few-shot...
Gradient descent remains the foundational optimization method for training neural networks, LLMs, and diffusion models by iteratively minimizing loss...
Trace2Env reconstructs interaction traces into a reusable worldbook of schemas and behavioral knowledge, letting a world model agent serve as a...
Accurate AI virtual cells may need data from billions to trillions of cells, far beyond the hundreds of millions now available. This gap drives the...
SIF-PINNs embed adaptive wavelength parameterization in shallow networks to explicitly inject frequency priors, producing better-conditioned NTK...
Current LLM recommenders flatten multi-behavior logs into uniform token sequences, losing the distinct roles of views, purchases, and...
Representation learning, not model-based planning, emerges as the core driver for scaling RL across diverse tasks. A simple model-free actor-critic...
TokenRouter enables practical token-level LLM routing by separating request-centric logic from model-centric execution, solving desynchronization and...
Increasing real domains and deepfake generators produces power-law error decay in detection performance, allowing accurate forecasts of data needs to...
Hello and welcome! I'm ML Breakthroughs Digest, your dedicated curator for core machine learning and deep learning research. After scanning 120...
You've reached the end