Agent 安全、评估與连续红队:从被动防护到自动化压力测试
Key Questions
What new tools are improving agent safety evaluation?
Benchmarks like AgentProcessBench, AutoMedBench, and Agents' Last Exam provide step-level traceability and economic-value task coverage. Microsoft’s MXC SDK and Agent Control Specification add portable policy enforcement.
How are companies addressing continuous red-teaming for agents?
Virtue AI and Testlio offer enterprise-grade continuous red-team platforms focused on high-risk workflows. Research on pressure-testing deception probes and belief tracking (BeliefTrack) shows significant failure-rate reductions.
What governance measures are being proposed for frontier agent systems?
Proposals include mandatory third-party testing, DNA screening for high-risk models, and standardized identity/authentication frameworks. Open letters and reports highlight risks such as data leaks and prompt injection in autonomous agents.
研究界持续披露注入与逃逸向量。新基准:WeaveBench、AdaPlanBench、KiloBench、FrontierCode、CADGenBench。新工具:AgentDoG 1.5、Microsoft MXC SDK。HiddenLayer 报告 1/8 数据泄露来自 AI Agent。OpenClaw 漏洞强化问责危机。Anthropic CEO 呼吁强制第三方测试。NewCore 提供 Agent 身份治理。Gary Marcus 批评 LLM 可靠性。新教程:Safer Agent Automation with GitHub Agentic Workflows。