Fine-tuning and agent ecosystem
RLSVR paper extends self-verifiable rewards to open-ended tasks with open-source code. Natolambert's RL train-inference parity fix for Qwen3.5 improves async RL. Unsloth Studio, Hermes Agent, OpenClaw, and ClawStack drive local agent development. Practical fine-tuning guides (LoRA, PEFT) and tools (MoFL) lower barriers. New: A guide for selecting open-weight instruct models for RL post-training was published, covering infrastructure and licensing.
Sources (3)
Updated Aug 8, 2026