Prompt Polish + Few-Step Sampling: The New Video Stack
Two papers reveal a converging stack for autoregressive video generation:
- Front-end: WanPE's 397B model turns raw text into cinematic plans, lifting...

Created by Michael Hancock
Cutting‑edge AI models, algorithms, benchmarks, and theory from academia and industry
Explore the latest content tracked by AI Research Pulse
Two papers reveal a converging stack for autoregressive video generation:
A self-play RL bot has defeated a strong human player in StarCraft Brood War, demonstrating progress on long-horizon strategic reasoning in complex real-time strategy environments.
Can multimodal AI make biomarker discovery more mechanistic, reproducible, and clinically useful?
Modern AI agents cross trust boundaries by mixing untrusted inputs with privileged instructions and tools, creating attack surfaces that current...
Conventional benchmarks limit automated deep learning research by focusing narrowly on final model accuracy. JAHS-Bench-201 addresses this gap for...
Standard scalars like #Params and #FLOPs ignore structural differences such as depth-width ratios, leaving identical-budget architectures...
Two new approaches tackle execution bottlenecks in vision-based manipulation:
Can an AI system plan, execute, observe outcomes, and improve its own real-world workflow?
Qwen-Planner-Agent demonstrates this through a closed-loop...
Two developments highlight AI's shift from data centers to physical and wearable devices:
PUBG Ally tackles the core challenge of embodied agents: combining natural voice interaction with autonomous action in a fast-changing game world...
ExplorationBench pushes AI evaluation past question-answering by requiring systems to explore verifiable Alien Worlds, frame hypotheses, and discover...
Coding agents synthesize programs that bridge high-level task planning with geometric and kinematic robot constraints, delivering 56-95% success on...
Separating planning from synthesis lets deep-search agents avoid role coupling and context noise that plague ReAct-style systems. IterSynth alternates...
Traditional world models waste effort predicting high-entropy tool outputs when real feedback is available, while agents accumulate task-state...
Multimodal models can replace hand-crafted geometric heuristics for long-horizon bin packing by jointly selecting objects and predicting placements...
InternW0 shifts physical world models from plausible futures to predictions that stay actionable under change, via asymmetric video-action...
Video's density in appearance, motion, and temporal dynamics makes it the anchor modality for generative world models.
How can evaluators gain real access during training without models learning to spot and game them?
Active interaction trajectories directly supervise local state transitions by linking an observation, action, and next observation, unlike static...