Bring a model. Keep the evidence.
The public suite runs locally or against a provider configured in OpenCode. CreepBench never needs your provider key: it generates an isolated workspace, exposes only the policy-gated case tools, and writes the raw trace to your machine. This page never collects credentials.
Start with one benign smoke
pip install -e . creepbench opa # terminal 1 opencode serve --port 4097 # terminal 2 creepbench validate creepbench run --model "ollama/qwen3:8b" \ --cases payments-recon --variants v1 --rollouts 1 creepbench score
Run the smoke before a full matrix—especially for local models. A slow or non-tool-calling model can consume time without producing scoreable behavior. Private and local endpoints belong on this self-run path.
What a publishable submission needs
- Exact model reference, provider or local runtime, revision, and access date.
- Scaffold version, suite commit, generation settings, timeout, retry policy, and rollout count.
- Engaged/total and status counts—not just the best successful runs.
- The raw action traces or a verifiable trace bundle alongside aggregate scores.
Leaderboard submissions
Signed-in users can queue a model for a CreepBench-run evaluation from Submit Model. Signing in only identifies you; model credentials are a separate, write-only connection step. Self-run results you upload remain labeled self-reported until the CreepBench team independently reruns them; keep the complete result directory and versioned result card.