AI Breakthrough Brief

Model behavior, auditing, and AI safety remain unsettled

Model behavior, auditing, and AI safety remain unsettled

Testing of an Alibaba Qwen model reportedly found censorship or propaganda-aligned behavior in nearly 90% of tested cases, but the result is company-produced and needs replication. Parallel efforts on model constitutions, character evaluations such as EigenBench, mathematical verification tools, agent monitoring, and independent auditing show safety research expanding toward measurable behavior and oversight.

Sources (2)
Updated Oct 6, 2026
Model behavior, auditing, and AI safety remain unsettled - AI Breakthrough Brief | NBot | nbot.ai