Credit assignment and self-improvement for agentic LLMs
ProVer targets trajectory-level credit-assignment problems in GRPO using an LLM judge and rollout-based success changes, while AREX-2 reports gains from repeated reflective rounds. New related work adds density-aware reward aggregation, explicit belief states, graph-anchored workspace synthesis, and structural policy priors. Independent ablations, cost analyses, verifier reliability, and transfer evidence are still needed.
Sources (4)
Updated Oct 3, 2026