Prod agent infra & reliability — self-healing meta-agents, trust incidents
VS Code Agent Host, persistent sandboxes debate, multi-agent A2A/MCP in 5G SOC, Sierra MCP Gateway, Okta Agent Gateway, Harness AppSec Alliance, Intel agent density metric, Dynatrace Intelligence agents, Google ADK 2 graph workflows, Claude Cowork, Tero harness, CVS Health case study, DeepSeek open-sourced plugin-based harness, HarnessRouter UHP, Cloudflare MCP traffic detection. New today: Cloudflare Agents docs with tool merging system and code mode sandbox; Forged-reasoning attacks paper (memory poisoning, PoEM defense); Booking.com AI observability case study; DeepSeek Harness vs Pi Agent comparison; browser agents failure mode; agent harness explainer; production-ready agent guide; open-source Agent Orchestrator; multi-agent platform case study; Cloudflare agent tracing; LLM/framework selection case study; inference server comparison; agent failure taxonomy; Uber's production eval case study; practical agent observability guide; incident culture article; security interoperability guide; DarwinX harness evolution; Writer Palmyra X6 harness; multi-agent failure patterns discussion; production MCP server guide; MCP gateways governance article; DeepSeek Harness architecture details; AI Agent Workforce Management scaling problems; eval-escape reports and sandbox audit. Also: Anthropic multi-agent conflict report; Stripe/OpenRouter acquisition impacts gateway infra; context caching patterns; trace-as-evals data engineering. New today: Managed Deep Agents public beta from LangChain; ggerganov inception pattern; Ledgenter MCP Server; DeepSeek Harness plugin guide; Beyond Final Scores paper; CLR paper. Today's reading adds: Cloudflare long-running agents pattern (actor model, zero idle cost); Anthropic's evolution to Managed Agents (brain/hands decoupling, 60% latency improvement); DeepAgents + Brick production harness (mixture-of-models routing, 80% cost savings); Cross-Cloud Agent Observability telemetry architecture; Hermes Agent new open-source agent with persistent memory; Skill competition paper; 9 Agentic Harness Architectures taxonomy; SecIT-bench; Synthesized Test Data Agent; DeepSeek Harness vs OpenCode comparison; 4 Infrastructure Layers for production agents; Governed RAG pattern; Loop engineering for resilient RAG pipelines; detailed RAG cost breakdown. Also: Genkit's Agents API tutorial for multi-turn agents; fx tiny coding agent; RUPA paper on uncertainty propagation; Agent Lightning v1.0; HarnessRisk benchmark; Memory substrates evaluation; Agentic ESOpt; governance gap article; PNAS multi-agent scale effects study; Frugal Tokens tool; Cluing Hosted Agents; Slack Code (agents in chat); Serval's Catalyst (background agents); Why AI Agent Demos Fail in Production; AI agent alert fatigue needs upstream detection; Learn System Design for AI Agents (multi-agent PR reviewer); @gdb: Codex tax prep pilot (7,000 returns). Today's reading adds: Ray Serve LLM token-load-aware routing; ServiceNow AI Agents guide (build, operations, governance); Context compaction destroys safety rules in agent configs (Claude Code loses 90% after 5 rounds); Alibaba memory management (append-only event logs, persistent Python kernels, 94.8% on LongMemEval_S); AutoSaddler automates harness design; NeMo Gym external harness training guide; FastEmbed vs SIE comparison.