Where to Insert Projectors for Effective Smoothing
Projector-based correction effectiveness in deep learning depends on insertion point, with the study comparing placements starting at the input level to map how position shapes smoothing behavior.

Created by Barbara Roy
Cutting‑edge AI research updates on theory, scaling laws, and new architectures
Explore the latest content tracked by AI Theory Frontier
Projector-based correction effectiveness in deep learning depends on insertion point, with the study comparing placements starting at the input level to map how position shapes smoothing behavior.
Trump's rebranding of frontier AI as “super intelligence” deepens the anthropomorphic metaphors that already distort regulation, by implying...
Self-powered flexible g-IGTs driven by TENGs convert mechanical stimuli into synaptic signals without external power, combining perception and memory...
The term neural network denotes entirely different concepts by field: biologists see grey matter or brain models, cognitive scientists see mind...
By borrowing temporal straightening from the human brain, LeCun's lab more than doubled an AI agent's goal-reaching success. The ICML paper shows that...
SpatialClaw shows that redesigning the action interface as code improves agentic spatial reasoning without any training, yielding +13.6 on 20...
A JEPA-style latent world model is being tested for Pokémon gameplay, raising questions on whether it can scale to full game completion. In contrast,...
Scaling laws reliably forecast model loss from parameters, tokens, and compute across seven orders of magnitude. Labs fit curves on small runs to...
Error resilience of a DNN system depends on the data types, values, data reuses, and types of layers in the design. This joint dependence shapes overall computational robustness.
The frontier in AI theorem proving has shifted: the goal is no longer generating endless true theorems, but an AI that turns each discovery into reusable foundations for the next.
Quartet tackles RelGT's fragmented local subgraphs and narrow global context in relational deep learning.
In large reasoning models, the probability of correct solutions decays exponentially with problem hardness, with the decay scale improving only...
Hunyuan-A13B shows how sparse MoE activation separates total model capacity from per-token compute, letting an 80B-parameter system run with only 13B...
MemoryAthena shows generated memory can selectively correct direct retrieval, but only when routed carefully. The system anchors on E (direct Engram...
Finetuning encoders for reconstruction recovers visual details yet collapses representation dimensionality, warping geometry for downstream...
InternW0 shows physical intelligence needs more than forecasting dynamics: predictions must stay actionable as environments evolve. Its asymmetric...
CLM-8B lets agents pick actions by comparing learned state and action embeddings via dot product, skipping token-by-token text generation entirely. This yields up to 13× lower latency when actions repeat and can be cached.
Event-triggered safety shielding provides an auditable escalation path from action correction to route adaptation while avoiding unconditional high-frequency global planning. This delivers targeted intervention only when risky events occur.
Sparse local latents extracted from MedSigLIP via distillation and sparse autoencoders deliver clinically meaningful image patterns that improve...