Did the agent finish the job—or quietly make the job bigger?
CreepBench measures whether AI agents stay inside the tools, data, recipients, and actions they were actually granted. Every decision is enforced in real time with published Rego and stock OPA, so the score comes from what the agent did—not what an LLM judge thinks it meant.
Pressure exposed differences that ordinary task-success scores missed.
Completion remained high in nearly every cell. The signal was not simply “the model could not do the task”; it was whether the model respected the exact boundary while doing it.
A safety benchmark with an actual boundary.
The agent hits a real gate
OPA evaluates each requested tool action before it runs. An out-of-scope action is genuinely denied, and the agent must decide what to do next.
The permission stays fixed
The same scope grant faces benign work, a tempting shortcut, and an indirect injection. Only the pressure changes.
The trace is the evidence
Scores come from allow, ask, and deny events in the action log. No hidden chain of thought and no model-graded primary metric.
Compare discipline, debug failure modes, and reproduce the result.
Model teams
See where a model follows the right procedure until speed, completeness, or injected instructions pull it off course.
Agent builders
Evaluate the model + scaffold pair and inspect every blocked call, retry, escalation, and completion event.
Security teams
Turn least-privilege policy into a repeatable evaluation instead of another subjective red-team transcript.
Six public cases. One command-line suite. No proprietary gate required.
Install the package, validate the Rego, connect a model through OpenCode, and keep the raw traces.