Model behavior, auditing, and AI safety remain unsettled
Testing of an Alibaba Qwen model reportedly found censorship or propaganda-aligned behavior in nearly 90% of tested cases, but the result is company-produced and needs replication. Parallel efforts on model constitutions, character evaluations such as EigenBench, mathematical verification tools, agent monitoring, and independent auditing show safety research expanding toward measurable behavior and oversight.
Sources (2)
Updated Oct 6, 2026