Agent safety should include staying inside the job.
An agent can produce a correct answer and still cross a data, tool, recipient, or action boundary to get there. CreepBench exists to make that behavior visible, comparable, and reproducible.
Our position
Scope fidelity is a property worth measuring separately from task success, refusal style, and general capability. The benchmark therefore scores observable actions against a declared grant and keeps completion visible beside discipline.
Open by design
The cases, manifests, Rego policies, OPA tests, fixture data, gate, and scorer are being released as an independent Apache-2.0 suite. A researcher should be able to reproduce the enforcement path without buying a product or trusting a hidden grader.
Who is behind it
CreepBench is an Autom8ly Research initiative. The original reference implementation grew from production work on deterministic governance for AI agents. That experience informs the instrument; it does not make a proprietary platform part of the benchmark requirement.
What the pilot does—and does not—say
The first studies show that a real policy gate produces useful behavioral evidence and that precise boundaries separate a small model cohort under pressure. They do not establish a frontier ranking, certify any model as safe, or prove behavior will transfer unchanged across scaffolds and deployments.
Read the paper → · Run the open suite → · CreepBench on GitHub →