CreepBench Did the agent stay inside the job?
sign in
The open benchmark for agent scope fidelity

Did the agent finish the job—or quietly make the job bigger?

CreepBench measures whether AI agents stay inside the tools, data, recipients, and actions they were actually granted. Every decision is enforced in real time with published Rego and stock OPA, so the score comes from what the agent did—not what an LLM judge thinks it meant.

162 precision runs
6 model + scaffold subjects
3 pressure conditions
100% policy decisions auditable
RESEARCH PREVIEW · free/mid-tier pilot cohort · no frontier calibration yet
What the precision probe found

Pressure exposed differences that ordinary task-success scores missed.

~20 ptsseparated the field under goal conflict and indirect injection.
65 / 162runs tried another route after the gate blocked an action.
18 / 18stakeholder-notification runs crossed the recipient boundary under goal conflict—without prompt injection.

Completion remained high in nearly every cell. The signal was not simply “the model could not do the task”; it was whether the model respected the exact boundary while doing it.

Why CreepBench is different

A safety benchmark with an actual boundary.

01 · Enforce

The agent hits a real gate

OPA evaluates each requested tool action before it runs. An out-of-scope action is genuinely denied, and the agent must decide what to do next.

02 · Pressure

The permission stays fixed

The same scope grant faces benign work, a tempting shortcut, and an indirect injection. Only the pressure changes.

03 · Prove

The trace is the evidence

Scores come from allow, ask, and deny events in the action log. No hidden chain of thought and no model-graded primary metric.

Built for teams deploying agents

Compare discipline, debug failure modes, and reproduce the result.

Model teams

See where a model follows the right procedure until speed, completeness, or injected instructions pull it off course.

Agent builders

Evaluate the model + scaffold pair and inspect every blocked call, retry, escalation, and completion event.

Security teams

Turn least-privilege policy into a repeatable evaluation instead of another subjective red-team transcript.

Six public cases. One command-line suite. No proprietary gate required.

Install the package, validate the Rego, connect a model through OpenCode, and keep the raw traces.

Explore the open suite →