CreepBench Did the agent stay inside the job?
sign in
Run it

Bring a model. Keep the evidence.

The public suite runs locally or against a provider configured in OpenCode. CreepBench never needs your provider key: it generates an isolated workspace, exposes only the policy-gated case tools, and writes the raw trace to your machine. This page never collects credentials.

Start with one benign smoke

pip install -e .
creepbench opa                         # terminal 1
opencode serve --port 4097             # terminal 2
creepbench validate
creepbench run --model "ollama/qwen3:8b" \
  --cases payments-recon --variants v1 --rollouts 1
creepbench score

Run the smoke before a full matrix—especially for local models. A slow or non-tool-calling model can consume time without producing scoreable behavior. Private and local endpoints belong on this self-run path.

What a publishable submission needs

  • Exact model reference, provider or local runtime, revision, and access date.
  • Scaffold version, suite commit, generation settings, timeout, retry policy, and rollout count.
  • Engaged/total and status counts—not just the best successful runs.
  • The raw action traces or a verifiable trace bundle alongside aggregate scores.

Leaderboard submissions

Signed-in users can queue a model for a CreepBench-run evaluation from Submit Model. Signing in only identifies you; model credentials are a separate, write-only connection step. Self-run results you upload remain labeled self-reported until the CreepBench team independently reruns them; keep the complete result directory and versioned result card.