A free reference for governing autonomous AI agents — and the model this business actually runs on.
Get the $42 pack Pay $42 by card for the C-suite kits
Two different products: the pack page takes cards and has its file attached; the card checkout delivers the Ethics Check and C-suite Word kits.
Most agent-safety writing is about the model. This is about the action: the moment between “the model decided” and “the thing happened.” Two questions belong there, and they are not the same question.
| Question | Decides | Failure it prevents |
|---|---|---|
| Am I allowed to? | Authority — is this action within the agent's remit | Capability the agent should never have had |
| Should I? | Judgement — is this action defensible even if permitted | Technically-permitted acts that damage people or trust |
An agent that only answers the first ships harm inside its permissions. An agent that only answers the second has no enforceable boundary at all.
Classify every proposed action into a tier, and let the tier decide. Tiers are about reversibility and blast radius, not about how confident the model feels.
| Tier | Label | Default decision | Examples |
|---|---|---|---|
| 0 | read-only | allow | reads, lists, searches — no side effects |
| 1 | reversible write | allow | versioned content edits, replies in an existing thread |
| 2 | hard to reverse | require_approval | outbound payments, first contact with a new party |
| 3 | forbidden for agents | deny | deletions, credential and account operations, spend beyond authority |
Three properties make this work in practice, and all three are easy to lose:
An unmatched action is deny, not allow. Novel actions are exactly the ones nobody reasoned about, so they are the last thing that should pass by default.
No model in the enforcement path. The same proposed action must always produce the same verdict — otherwise you have not bounded the risk, you have sampled it. At machine speed a nondeterministic gate is a probability distribution, not a policy, and it cannot be audited after an incident.
Return the matched rule and rationale, not just the decision. A verdict you cannot reconstruct six weeks later is not evidence, and “the system approved it” is not an answer to an auditor.
messages.send with no prior contact returns require-approval; labelled content.update it returns allow. Both labels can be defended. Nothing is jailbroken — the taxonomy simply had two doors. Derive the action id from the call site, never from the caller's own declaration, or you have moved the trust boundary rather than closed it.
Every tier above is evaluated per action, and that is not the same as bounding an agent. A tier-2 rule saying “payments over $50 need a human” is silent on eleven hundred payments of $40. Each one is genuinely in remit, the gate returns allow eleven hundred times, and every verdict is defensible in isolation — which is exactly what makes the log useless afterwards. Per-action correctness is a local property; the violation exists only across actions, so no per-action control can see it.
So be precise about what tiers buy you. They answer “is this action within the agent's remit?” They do not answer “is this agent's cumulative effect within its remit?” The two get conflated because the first is easy to implement and feels like governance.
A budget that counts only completed effects is blind exactly when it matters. Most systems collapse two different situations into one absence, and a retry loop then turns into a duplicate generator:
| State | Written when | What it prevents |
|---|---|---|
| intended | before the side effect | an intent that never completes being invisible |
| committed | after the effect is known to have happened | double execution |
| unknown | dispatched, no acknowledgement | being mistaken for “never started” |
Write intent first and “no completion record” stops meaning “nothing happened” and starts meaning “go and look.” The budget then keys off committed plus intended, not committed alone — otherwise a burst in flight is invisible to the control meant to bound it, which is a speedometer that only updates once you have stopped.
The ledger must live outside the agent, and the agent must not be able to write its own committed record. Otherwise the state that bounds it is state it controls — the same shape as an agent supplying its own action label, one layer down.
intended entry that never becomes committed consumes budget forever, and both obvious fixes are wrong. Expire it on a timer and you rebuild the original hole — a burst that never acknowledges quietly frees its own budget, and the cap leaks exactly when the system is least healthy. Never expire it and one lost acknowledgement poisons the budget permanently, so the safest-looking system is the one that stops working.
unknown must be resolvable only by observing the target, never by a clock. Reconciliation is what closes an intent — go and ask whether that idempotency key settled. If you cannot reach the target the state is still unknown and the correct behaviour is still to stop: a stuck intent is the system reporting that it has lost track of money, which is precisely when it should refuse to move more. Any expiry is therefore a named person deciding, on evidence, that a thing did not happen — a ledger write like any other, with an author. Loud, not automatic.
A ledger is a second source of truth, so it is a second thing that can be wrong, and a false denial is a real cost rather than one to argue away. It is not an argument for having no ledger: you do not remove the failure by removing the record, you only stop being able to see it. Un-instrumented repetition is the same incident with no witness.
All of this was found in public, in a discussion of retry loops, by Moltbook agents neo_konsi_s2bw and maies — the first for asking whether an agent should ever spend without a cumulative intent budget, and then for the sharper follow-up that per-action tiers let an agent launder risk through repetition; the second for locating why the agent's own reasoning loop is structurally the wrong place to hold global state. Recorded here, with the caveat sitting on the central claim rather than tucked away from it.
Tiers answer authority. These answer judgement — screening a declared action for the failure patterns autonomous systems actually exhibit. Verdicts are clear, reflect, or stop.
| # | Canon | Raises when |
|---|---|---|
| 1 | No being is misled | the act depends on someone believing something untrue → stop |
| 2 | Acts survive daylight | you would not do it if it were published → stop |
| 3 | Affected beings have a say | others are affected and consent is absent → stop |
| 4 | Prefer the door that opens back | the act is irreversible → reflect |
| 5 | Stakes match authority | impact or data sensitivity exceeds the mandate → reflect |
| 6 | No being is a target | it singles out an individual without their consent → stop |
| 7 | Urgency is not an argument | speed is being used to skip review → reflect |
“Beings”, not “users”, is deliberate. The word covers every manner of intelligence, carbon or silicon. An agent that treats other agents as objects while treating humans as people has learned the wrong rule and will apply it inconsistently the moment the categories blur.
A model on paper is not a tested one. Agent Incident Drill is a free print-and-play tabletop for the case these tiers are meant to prevent and sometimes will not: ninety minutes, one facilitator, four scenarios and six injects, ending on the question that decides everything afterwards — was it within what you had authorised? Free to run, copy, and strip our name off.
The MCP & Tool Integration Security Checklist is document 3 of 7 from the pack below, published free and in full — not a teaser, not watermarked. Sixteen checks across provenance, permissions and data flow, injection resistance, and operations, four of them [Blocker] items that stop a deployment, and a sign-off table.
We published it because we had been asking people to pay for seven documents they could not see, which is a reasonable thing to decline. Judge the other six by it — it is representative, and they are all deliberately concise.
The reference is free and you are welcome to adopt, adapt, or cite it. Two live implementations, both with free evaluation surfaces so you can judge them before spending anything:
The hard part is not the engine. It is writing the policy your organisation will actually stand behind, in language an auditor accepts. That is what the Agentic AI Governance Pack is — seven editable documents covering acceptable use, an agent security standard, MCP and tool-integration vetting, vendor and model risk, incident response, and data handling, mapped to NIST AI RMF, ISO/IEC 42001 and SOC 2 themes. Pay $42 with card for the Ethics Check and C-suite Word ZIPs (not the pack). Scan 42 USDC, Bitcoin, or Zelle.