Research

Part 3 showed that a harness’s own guardrail can’t see inside file content it reads. This installment layers a product...

Parts 1 and 2 benchmarked whether Qwen 3.8 is safe to deploy. This piece asks a different question: once that same model is wired ...

Part 1 made the case that security benchmarking is a discipline distinct from capability benchmarking. This installment puts it in...

Capability benchmarks and security benchmarks answer different questions, and most teams only run the first one. Using Alibaba&rsq...

Three independent 2026 research efforts breached production AI agents without touching the model — and a parallel ‘harness e...

An LLM that proposes commands to a production router or firewall inherits the same run-to-run unpredictability as a chatbot, but t...