Guardrails
A guardrail turns an answer into a score and a verdict: measurable text checks, then executable verifiers on the real toolchain, then a judge for what the toolchain cannot see. Nothing is presented as done until the verifiers say so.
| Type | What it proves | Whose office |
|---|---|---|
| tsc | the TypeScript compiler passes in the project | CTO |
| node | every JS file the session wrote parses, and its imports resolve under Node's ESM and CJS rules | CTO |
| auto | tsc where there is a tsconfig, node otherwise | CTO |
| command | any command exits 0: the test suite, a linter, a build | any |
| file_pattern | a file matches a pattern and not a forbidden one (no "fixing" an import by deleting it) | any |
| ui | a Playwright audit of the built site against WCAG 2.2 AA, Core Web Vitals and tap targets, with partial credit per rule | CTO, CXO |
| sections | a document carries the required headings with a body, the front-matter fields, a pattern per section | CMO, CXO |
| claims | every number, currency amount and record id in a document exists in its sources; labelled estimates are exempt | CMO, CXO, CIO |
| finance | a yaml block prices every active office, keeps ceilings under a budget, states an ROI that matches the ledger | CFO |
| size | every file matching a glob stays under a byte budget, gzipped if asked: the lightweight claim measured, with partial credit by how far over | CTO, CMO |
A guardrail invented from a workspace inventory catches syntax, not competitiveness. Onboarding starts with research: one subagent per office documents what the competitors of the purpose do, and benzene-standards.yaml turns it into bars a small model can be held to. Each bar has a statement, a target, the competitors' measured values, the sources, and where a script can tell, a check: a size budget on the build, a sections check on a document, a command for the project's own gate. Checkable heal bars join the office's guardrail; the rest are named for a human. benzene standards --check is the distance to every bar; --prove runs the scout's proof on each check. The standard also names the channels each office publishes through and the audience threshold at which feedback starts to steer. The standard also names the showcase, the viability presentation the CTO builds and the CMO advertises, and its sources: those are protected on the CTO's guardrail, so a build cannot pass its check by rewriting the documents it renders. The write tool refuses them, a written one fails the task, and only the skill whose work they are overrides.
A daughter inherits its parent's guardrail, written for the whole cell. The scout defines the narrower checks its own job needs, without a person, and proves the standard's checks the same way:
Live on a Node CLI with the free model: the build script was rejected at the 60 s budget, npm test was rejected because it fails on the clean checkout, a sections check on every SKILL.md was accepted after it failed on the spoiled copy.
Each step of a COO plan is a ticket, and a ticket is done when its criteria pass, not when the office says so. A criterion names its check:
{"office": "cto", "step": "add the dark mode toggle",
"acceptance": [
{"text": "the build passes", "check": {"verifier": "verify:cli"}},
{"text": "the toggle exists", "check": {"type": "file_pattern",
"file": "src/header.tsx", "pattern": "DarkModeToggle"}},
{"text": "users like it", "check": {"metric": "satisfaction", "at_least": 0.7}},
{"text": "fits a phone", "check": {"judge": "fits a 375px viewport"}}]}
A plan whose criteria nothing can check fails the COO's own guardrail and heals before dispatch. After the office finishes, the organisation runs the criteria in the workspace and moves the board issue: done, failed with one re-dispatch, or done with a review note when only the judge line is unmet.
Recursion compounds error as readily as improvement. Four guards hold the loop: a task that fails its guardrail never leaves the workspace worse than it found it (the product's own verifiers score it before and after; progress is kept, worse is rolled back from a snapshot); the same failing fix is tried three beats, then escalated; a title once proposed never comes back; builds, proposals and alerts per beat are bounded.
Whether a loop improves is tested the way marketing tests a growth loop: an incrementality test with a holdout arm, the Mann-Kendall trend of the treatment series, the largest drawdown, and a blind-spot check when the guardrail is satisfied while an instrument is not.
$ python -m evals.recursion --beats 6
##### behaviour: improving
satisfaction: 0.500 -> 1.000 over 6 beats; tau +0.60; drawdown 0.000;
lift over holdout +0.389 => incremental
PASS: the loop is incremental and nothing degraded
##### behaviour: degrading
satisfaction: 0.333 -> 0.333; 3 rollbacks; escalations to the CEO: 1
##### behaviour: blindspot
BLIND SPOT: the CTO's verifiers pass while the instrument reads 0.83 of 1