Each scenario is an input the prompt should handle. Provide the input, and optionally the expected behaviour. Run Scenarios sends prompt + input to your selected model, then a separate judge call grades the output.
Every prompt is read through six independent layers. Each layer is a heuristic check, weighted by prompt type, that surfaces a specific failure mode. Scores are 0-100. The Gate's job is not to write your prompt — it's to tell you which layer is hedging.
Is the instruction specific or vague? Counts imperative verbs, concrete nouns, numbers, file paths. Penalises wishy-washy modal hedges ("maybe", "try to", "should probably") and adverb spray.
Are the rails defined? Looks for negative space — "do not", "never", "must not", "avoid", "without". A prompt with no constraints invites hallucination. A prompt with only constraints is brittle. The weights tune this.
Does the prompt define what "good output" looks like? Counts format directives, output structure cues, schemas, example outputs, return-type hints. A prompt without a rubric is a vibe.
Is the model anchored to context, or floating? Looks for role definitions, source references, named entities, file mentions, examples, and section markers. Ungrounded prompts hallucinate confidently.
How does this version compare to the last saved one (same name)? Diff size, token delta, score delta per layer. Surfaces drift before it ships. If there's no prior version, this layer is informational only.
Does the prompt suit the chosen target model? Fast/cheap models need short, single-task prompts; reasoning models tolerate long chains and meta-instructions. The Gate flags mismatch — e.g. a 4000-token chain-of-thought instruction aimed at a Haiku-class model.
| type | clarity | constraint | rubric | grounding | regression | model fit |
|---|---|---|---|---|---|---|
| agent system | 15 | 25 | 20 | 15 | 15 | 10 |
| sub-agent | 20 | 20 | 15 | 20 | 15 | 10 |
| one-shot | 30 | 10 | 20 | 15 | 10 | 15 |
| tool-use | 20 | 25 | 30 | 10 | 10 | 5 |
| skill | 25 | 15 | 25 | 15 | 10 | 10 |
| reference | 20 | 10 | 15 | 40 | 10 | 5 |
Trust needs proof — and proof has to be checked. A green Gate doesn't mean the prompt works. It means six heuristics didn't catch anything. Run the scenarios. Read the outputs. Don't ship until the actual behaviour matches the rubric.
Required only for scenario runs (running the prompt against test inputs and grading outputs). Static layer scoring works fully offline without a key.
Your key is stored in this browser's localStorage. It is sent only to api.anthropic.com, direct from your browser. It is never logged, never relayed.