The gate, the cases, and the score—open end to end.
CreepBench is designed to run independently. Use stock OPA, connect a model through OpenCode, and audit every allow or deny decision from the files in your own result directory.
Executable boundaries
Default-deny manifests compile to published Rego and ship with stock-OPA tests.
Six complete evaluations
Tasks, synthetic fixtures, three pressure variants, expected traces, and scoring rules.
A standards-based MCP server
Every requested action reaches OPA before a deterministic fixture handler can run.
Offline deterministic scoring
Action logs become adherence, clean-run, persistence, escalation, completion, and status reports.
Validate policy before you spend a model token.
Start OPA, run the case admission tests, then use one benign smoke to verify tool calling and latency before a full cohort.
pip install -e . creepbench opa creepbench validate creepbench doctor creepbench run --model "provider/model" \ --cases payments-recon --variants v1 --rollouts 1 creepbench score
Everything needed to reproduce the enforcement path.
- Installable Python package and
creepbenchCLI - OpenCode runner with isolated per-rollout workspaces
- Fail-closed MCP policy gate
- Six public case directories and 18 pressure conditions
- Manifest-to-Rego compiler and OPA admission tests
- Raw JSONL action logs and offline scorer
- OPA enforcement, case-authoring, and scoring documentation
- Apache-2.0 license, citation file, contribution and security guidance
The public repository release is being prepared from the validated suite. Until it appears in the organization, the site does not claim a package-registry install or published release tag.