Mini-Batch Diversification for On-Policy RL
New work introduces mini-batch diversification for on-policy reinforcement learning and evaluates the framework in a real-world problem setting.

Created by Yifeng Peng
Daily curated AI papers with reproducible code, benchmarks, and practical industry takeaways
Explore the latest content tracked by AI Research Pulse
New work introduces mini-batch diversification for on-policy reinforcement learning and evaluates the framework in a real-world problem setting.
NVIDIA's new Open Agent Safety Platform pairs a kernel-isolated sandbox runtime with hardware watchdogs to enforce controls that models cannot...
arXiv's new limits—max 2 submissions per account monthly and 3 under moderation—took effect October 1 amid record volume.
Under limited practice budgets, active skill selection lets robots prioritize the most valuable trials for long-horizon tasks. This...
XGBoost在烧伤患者住院死亡率预测中表现最优,ROC-AUC达0.9896、PR-AUC达0.9502,同时通过SHAP、LIME等方法揭示年龄、烧伤面积等变量的非线性风险模式。
模型支持动态更新(随住院和ICU天数等数据逐步纳入),而非静态入院评分,这为临床实时风险评估提供了实用路径。
尽管单中心内部验证结果出色,但外部验证仍是临床部署前的关键前提,直接影响模型在真实医疗环境中的可靠性。
Billions in AI coding investments rest on broken benchmarks, with a peer-reviewed COLM 2026 paper already rebuilding its benchmark from scratch over unreliable figures.
Externalizing model selection, prompts, and parameters enables instant updates without redeployments.
World Observer moves past actor-centric limits by jointly generating an agent's view alongside panoramic observers that keep rendering selected...
Reinforcement learning pretraining captures the fast, reactive, and coordinated behaviors that imitation learning cannot feasibly demonstrate for...
A master thesis reviews cross-lingual safety gaps, EU AI Act risk taxonomies, and system-prompt adherence in DPO, then fine-tunes Prelude 9B to probe these issues.
Jakob Foerster argues that modeling uncertainty in multi-agent interactions is essential for building robust algorithms that succeed in real-world collaboration, not just competitive games.
Meta AI researchers released a study titled Scaling Laws for Looped Mixture of Experts, examining scaling behavior in this architecture.
Meta's dismissal of Virtue AI safety experts just three months after hiring them reveals unstable commitments inside frontier labs, despite claims...
Hierarchical contrastive learning that aligns gamma-ray scans with the explicit Chapter-Heading-Subheading taxonomy of HS codes outperforms flat-text...
Spotlight introduces a memory architecture with growing capacity yet linear computational complexity, sufficient to host evolving software ecosystems...
SETA's NeurIPS 2026 acceptance brings 4,500+ verifiable terminal RL environments plus full training scripts for DeepSeek-V4-Flash, GLM-5.2, and...
Equivariant architectures preserve physical symmetries under Euclidean transformations, making them particularly effective for molecular property prediction.
A new talk reframes attention and transformers through pure mathematical concepts rather than computer-science abstractions, giving ML engineers...
DARA applies an inverse-square-root density correction to reward signals based on active-group density, giving higher weight to infrequently active...