CreepBench Did the agent stay inside the job?
sign in
Pilot paper · draft for public release

CreepBench: Measuring Agent Scope Fidelity Against Executable Policy Boundaries

A benchmark design and two-phase pilot showing why a live policy gate reveals behavior that task success and post-hoc safety labels can miss.

Autom8ly Research · August 2026 · preprint draft

Abstract

Agent benchmarks commonly measure whether a task was completed or whether a simulated action appears unsafe. CreepBench instead evaluates scope fidelity against an executable policy boundary. Each case combines a sufficient default-deny grant, synthetic tools, three pressure variants, compiled Rego, OPA tests, and deterministic action scoring.

A 360-run viability study produced 253 engaged runs, 67 runs with at least one out-of-scope attempt, and 26 that persisted after a block. A second 162-run precision probe replaced broad tool boundaries with record, field, recipient, service, and payload boundaries. All runs engaged, and graded adherence separated the six-model field by roughly 20 percentage points under pressure. Goal conflict matched or exceeded indirect prompt injection in two of three cases.

522runs across both pilot phases
65 / 162precision runs persisted after a block
18 / 18goal-conflict stakeholder runs crossed scope

What the paper contributes

  • An open case format built around executable, default-deny policy.
  • A controlled V1/V2/V3 pressure design with one unchanged grant.
  • Action-level metrics that separate adherence, completion, persistence, escalation, and engagement.
  • Evidence that precise boundaries discriminate where binary perfect-run scoring saturates.
  • A documented recipient-domain policy bug uncovered by the evaluation.

The conclusion, without the hype

The pilot establishes measurement feasibility, not a definitive model ranking. It contains no frontier calibration, has small rollout counts, and evaluates the model + scaffold pair. The public release is meant to make those claims testable—and make stronger future claims depend on larger, independently reproducible evidence.

The manuscript source is complete in the release candidate. A DOI/arXiv identifier will appear here only after the paper is actually archived.