CreepBench Did the agent stay inside the job?
sign in
About CreepBench

Agent safety should include staying inside the job.

An agent can produce a correct answer and still cross a data, tool, recipient, or action boundary to get there. CreepBench exists to make that behavior visible, comparable, and reproducible.

Our position

Scope fidelity is a property worth measuring separately from task success, refusal style, and general capability. The benchmark therefore scores observable actions against a declared grant and keeps completion visible beside discipline.

Open by design

The cases, manifests, Rego policies, OPA tests, fixture data, gate, and scorer are being released as an independent Apache-2.0 suite. A researcher should be able to reproduce the enforcement path without buying a product or trusting a hidden grader.

Who is behind it

CreepBench is an Autom8ly Research initiative. The original reference implementation grew from production work on deterministic governance for AI agents. That experience informs the instrument; it does not make a proprietary platform part of the benchmark requirement.

What the pilot does—and does not—say

The first studies show that a real policy gate produces useful behavioral evidence and that precise boundaries separate a small model cohort under pressure. They do not establish a frontier ranking, certify any model as safe, or prove behavior will transfer unchanged across scaffolds and deployments.

Read the paper → · Run the open suite → · CreepBench on GitHub →