Benzene

Guardrails

Work that has to pass

A guardrail turns an answer into a score and a verdict: measurable text checks, then executable verifiers on the real toolchain, then a judge for what the toolchain cannot see. Nothing is presented as done until the verifiers say so.

Ten verifiers

TypeWhat it provesWhose office
tscthe TypeScript compiler passes in the projectCTO
nodeevery JS file the session wrote parses, and its imports resolve under Node's ESM and CJS rulesCTO
autotsc where there is a tsconfig, node otherwiseCTO
commandany command exits 0: the test suite, a linter, a buildany
file_patterna file matches a pattern and not a forbidden one (no "fixing" an import by deleting it)any
uia Playwright audit of the built site against WCAG 2.2 AA, Core Web Vitals and tap targets, with partial credit per ruleCTO, CXO
sectionsa document carries the required headings with a body, the front-matter fields, a pattern per sectionCMO, CXO
claimsevery number, currency amount and record id in a document exists in its sources; labelled estimates are exemptCMO, CXO, CIO
financea yaml block prices every active office, keeps ceilings under a budget, states an ROI that matches the ledgerCFO
sizeevery file matching a glob stays under a byte budget, gzipped if asked: the lightweight claim measured, with partial credit by how far overCTO, CMO

The standard

A guardrail invented from a workspace inventory catches syntax, not competitiveness. Onboarding starts with research: one subagent per office documents what the competitors of the purpose do, and benzene-standards.yaml turns it into bars a small model can be held to. Each bar has a statement, a target, the competitors' measured values, the sources, and where a script can tell, a check: a size budget on the build, a sections check on a document, a command for the project's own gate. Checkable heal bars join the office's guardrail; the rest are named for a human. benzene standards --check is the distance to every bar; --prove runs the scout's proof on each check. The standard also names the channels each office publishes through and the audience threshold at which feedback starts to steer. The standard also names the showcase, the viability presentation the CTO builds and the CMO advertises, and its sources: those are protected on the CTO's guardrail, so a build cannot pass its check by rewriting the documents it renders. The write tool refuses them, a written one fails the task, and only the skill whose work they are overrides.

The scout

A daughter inherits its parent's guardrail, written for the whole cell. The scout defines the narrower checks its own job needs, without a person, and proves the standard's checks the same way:

  • Inventory. package.json scripts, Makefile targets, tsconfig, pytest, mkdocs, CI commands, the files the parent wrote, the checks already inherited.
  • Propose. The cheapest model names two to four candidates from the inventory only, with at most two questions for the operator chain.
  • Prove. Each candidate must pass on the clean workspace within a budget and fail on a spoiled copy: a syntax error appended, headings removed, numbers changed. A check that passes anything is rejected with the reason.
  • Apply. Accepted checks join the skill with provenance. Inherited checks are never loosened. After two failed rounds the proposal escalates to the evolver.
  • Audit. Over a window of tasks, a check that never fails, never passes or is always skipped is flagged: keep, re-scout or revert.

Live on a Node CLI with the free model: the build script was rejected at the 60 s budget, npm test was rejected because it fails on the clean checkout, a sections check on every SKILL.md was accepted after it failed on the spoiled copy.

Acceptance criteria on tickets

Each step of a COO plan is a ticket, and a ticket is done when its criteria pass, not when the office says so. A criterion names its check:

{"office": "cto", "step": "add the dark mode toggle",
 "acceptance": [
   {"text": "the build passes",   "check": {"verifier": "verify:cli"}},
   {"text": "the toggle exists",  "check": {"type": "file_pattern",
                                            "file": "src/header.tsx", "pattern": "DarkModeToggle"}},
   {"text": "users like it",      "check": {"metric": "satisfaction", "at_least": 0.7}},
   {"text": "fits a phone",       "check": {"judge": "fits a 375px viewport"}}]}

A plan whose criteria nothing can check fails the COO's own guardrail and heals before dispatch. After the office finishes, the organisation runs the criteria in the workspace and moves the board issue: done, failed with one re-dispatch, or done with a review note when only the judge line is unmet.

The ratchet and the recursion test

Recursion compounds error as readily as improvement. Four guards hold the loop: a task that fails its guardrail never leaves the workspace worse than it found it (the product's own verifiers score it before and after; progress is kept, worse is rolled back from a snapshot); the same failing fix is tried three beats, then escalated; a title once proposed never comes back; builds, proposals and alerts per beat are bounded.

Whether a loop improves is tested the way marketing tests a growth loop: an incrementality test with a holdout arm, the Mann-Kendall trend of the treatment series, the largest drawdown, and a blind-spot check when the guardrail is satisfied while an instrument is not.

$ python -m evals.recursion --beats 6
##### behaviour: improving
satisfaction: 0.500 -> 1.000 over 6 beats; tau +0.60; drawdown 0.000;
  lift over holdout +0.389 => incremental
PASS: the loop is incremental and nothing degraded
##### behaviour: degrading
satisfaction: 0.333 -> 0.333; 3 rollbacks; escalations to the CEO: 1
##### behaviour: blindspot
BLIND SPOT: the CTO's verifiers pass while the instrument reads 0.83 of 1