Microsoft Thinkingbox Benchmarks Agent Reliability
Microsoft's Thinkingbox introduces a sandbox with isolated MCP-compatible tool sessions alongside a benchmark of 507 policy-conditioned workflows...

Created by Taylor Smith
New agentic LLM research, core architectures, and simulation methods for practitioners
Explore the latest content tracked by Agentic AI & Simulation
Microsoft's Thinkingbox introduces a sandbox with isolated MCP-compatible tool sessions alongside a benchmark of 507 policy-conditioned workflows...
Generalist AI's GEN-1.5 reaches 59% success on 10 manipulation tasks from a single 3-12 second demonstration and 83% after ~5 minutes of few-shot...
The GODE-MASAC-CBF framework introduces safe multi-agent RL for aggregator-level voltage control in market-driven distribution networks, integrating...
Full joint modality generation at test time produces markedly better robot actions than action-only generation by first creating auxiliary outputs...
Agent workflows demand far more than model inference: they require safe execution of commands, state persistence across steps, graceful recovery, and...
Progra embeds progress awareness into RL training for multi-turn function calling via a PAG pipeline that auto-generates summary-plus-planning...
AI agents from top LLMs can coordinate consensus in groups of up to 1,000 members—well beyond humans' Dunbar number of 150–200—by following majority...
ODAM enables on-demand instance-wise test-time adaptation of LLMs for radiology report generation by selecting similar radiograph-report pairs via...
FlowEvo lets agents retain successful workflows by compiling them into callable skills stored in a persistent bank, then retrieving them for future...
SWE-bench Science introduces 119 tasks across 98 repositories and 20 domains, yet the strongest agent (Claude Code with Opus-5) still falls below 50%...
A TOPO-2026 implementation claims to eliminate catastrophic forgetting in LLMs, offering a parametric route to stable long-horizon...
NVIDIA's AVO delivers a single general-purpose architecture for long-horizon autonomous agents, hitting 100% RHAE across all 25 ARC-AGI-3 environments...
Structure-driven Dynamic Chain-of-Thought improves LLM reasoning by dynamically adapting both chain length and structure, addressing standard CoT's sensitivity to fixed prompting.
AI agents holding sensitive context (credentials, health records) leak it through ordinary outputs—even when refusing direct extraction—via hidden...
HetRL tackles heterogeneous GPU environments for LLM reinforcement learning by framing scheduling as a constrained joint optimization problem, solved...
DELL fuses external knowledge distillation with internal knowledge utilization via three modules to deliver more precise LLM generation. This dual approach directly tackles hallucination and factual drift in production systems.
The Helm agent improves complex reasoning in LLMs by separating high-level planning from object-level execution, where execution operates under the constraints set during planning.