Installing a governor beside an autonomous coding agent creates a second system that can make consequential decisions. If the governor can block tools, it can also block the exact verification, read or retry that would have completed the task. That means “turn the guardrail on” is not a neutral configuration change. It is a transfer of authority.

In short

  • Installation is not evidence. A new governor has no local track record for this repository, engine version, workflow or action family.
  • Observation and authority should be separate. Shadow Mode can collect falsifiable evidence without changing the agent's next action.
  • Trust should be narrow and reversible. Evidence from one engine or repository should not silently authorize another, and drift should remove authority.
  • Fail-open is contextual, not universal. It is a conservative choice for a no-progress efficiency governor; security authorization controls may rationally choose different failure semantics.

The governor is another source of failure

Agent governance is often drawn as a simple control relationship: user, agent, guardrail. That picture hides the fact that the guardrail has its own implementation bugs, blind spots and assumptions.

USER ↓ AGENT ↓ GOVERNOR The governor can now be wrong too.

A runtime governor can misclassify useful verification as waste. It can receive incomplete hooks after an engine update. It can normalize two distinct actions into the same identity. It can believe state is unchanged because its fingerprint missed the relevant part of a large workspace. It can be perfectly correct on one integration and unsafe on another.

Once a control plane can deny work, those errors are no longer reporting errors. They are behavioral changes to the agent.

Installation should not equal authority

Most permission systems assume the operator already knows what the software should be allowed to do. Agent governors have a different problem: the thing they need to know—whether their intervention policy is safe and useful in this local workload—often cannot be established before observation.

At install time, the governor usually has no evidence about:

  • how this repository uses repeated reads or tests;
  • which lifecycle events the current engine version reliably exposes;
  • whether workspace-state evidence covers the relevant changes;
  • how often a recommendation would have been a harmful false stop;
  • the runtime overhead introduced by the governor itself.

Installation proves only that the software is present. It does not prove that intervention is justified.

Shadow Mode turns predictions into testable claims

The simplest way to separate observation from authority is Shadow Mode. The governor evaluates the proposed action, records what it would recommend, but does not change what the agent actually does.

agent proposes action ↓ governor evaluates evidence ↓ agent action is still allowed ↓ reality reveals whether the recommendation was justified

This is more than a “safe default.” It creates counterfactual evidence. If the governor predicts that another identical read will produce no progress, Shadow Mode lets the read happen. The outcome can then weaken or strengthen the rule.

Did workspace state change? Did new evidence appear? Was the action actually successful? Did a verifier later show that the repetition mattered? A recommendation that cannot survive those questions should not become authority.

Earned enforcement should be local and narrow

Once enough local evidence exists, the control problem changes. The question is no longer “is this policy enabled?” but “under exactly which conditions has this policy earned the right to interfere?”

SHADOW ↓ representative local observations ↓ coverage + outcome quality + false-stop review ↓ explicit promotion ↓ LIMITED ENFORCEMENT ↓ drift / weak evidence / integrity failure ↓ SHADOW

MARGINAL calls this Earned Enforcement. The important part is not the name. The important part is that authority is bound to evidence and conditions rather than to the fact that a package was installed.

In the current project, Codex can reach a narrow Tool Enforcement boundary for eligible local actions after repository-local evidence and explicit promotion. That is deliberately not described as Full Compute Enforcement. Claude Code and OpenCode remain Observe-only because their current surfaces and evidence do not justify broader claims.

Read the current capability labels and integration boundaries →

Trust should not silently transfer

Suppose a governor behaves well for Codex in repository A. That does not establish safety for Claude Code in repository B. Even if both integrations ultimately map events into the same provider-neutral protocol, they may differ in outcome fidelity, tool semantics, lifecycle coverage and actual ability to intercept actions.

EvidenceWhat it can supportWhat it cannot prove
Good behavior in one repositoryLocal confidence under the measured workloadSafety across unrelated repositories
Reliable outcome hooks on one engineStronger evidence for that engine/versionEquivalent outcomes on another engine
Ability to block one tool familyTool Enforcement for that supported boundaryFull control of model turns, hosted tools or retries

Provider-neutral policy is valuable only if it does not erase provider-specific uncertainty.

A trustworthy governor needs demotion, not just promotion

Traditional permission systems are good at granting access and bad at taking it back automatically. Agent runtimes change too quickly for that assumption. An engine update can alter hook payloads. A repository can change shape. Coverage can become incomplete. An identity or evidence hash can drift.

A governor that was justified yesterday should be able to say:

I earned authority under condition X. Condition X is no longer true. Therefore my authority is no longer justified.

This is why MARGINAL treats trust as reversible. Drift, unknown outcomes, integrity problems or weakened evidence can demote an integration back toward advisory behavior instead of preserving authority through inertia.

Fail-open is a contextual choice, not a universal safety law

“Fail open” can sound reckless in security engineering. Sometimes it is. An authentication or authorization boundary may rationally choose fail-closed behavior because allowing an unauthorized action is the dominant risk.

MARGINAL's fail-open argument is narrower: for a no-progress efficiency governor, uncertainty should not automatically become permission to block useful agent work.

The error costs are asymmetric:

FailureTypical effect for a no-progress governor
False negativeA redundant action may run and consume some additional time or compute.
False positiveA useful action is blocked; the task may fail, require human recovery and cause the governor to be disabled entirely.

That does not prove fail-open is always correct. It explains why ambiguity in this specific control problem should weaken intervention pressure rather than increase it.

The governor has to pay its own tax

A governance layer can make the total system worse even when it successfully reduces agent activity. It adds code paths, state, latency and possibly additional model calls. If evaluation counts only the agent compute it avoided and ignores the control-plane cost, the economics are incomplete.

net governance value = avoided low-value work - governance latency - governance compute - harmful interventions - recovery cost

This is also why MARGINAL does not promote its historical exploratory 24.93% token difference into a causal savings claim. In that smoke, neither lane resolved a task and no deny was applied. The project publishes the observation while refusing the attribution.

Inspect the public benchmark report and its limitations →

“Guardrail” is too broad a capability label

A system that can observe telemetry, a system that can recommend an action and a system that can actually intercept a tool call are not equivalent. Neither is a system that can stop local tools equivalent to one that controls model turns, retries, hosted execution and compute accounting.

MARGINAL's contributor-facing capability model separates:

  • Observe — telemetry and non-blocking recommendations;
  • Tool Enforcement — supported tool actions can be blocked or changed;
  • Full Compute Enforcement — model turns, tools, retries and stop behavior are controllable and measured.

The names are less important than the discipline: capability claims should describe the interception boundary that actually exists.

The principle: evidence before authority

An AI agent governor should start powerless for the same reason a scientific claim should start unproven: the burden belongs to the mechanism asking for trust.

Observe first. Collect local outcomes. Measure false stops. Verify the integration boundary. Account for the governor's own cost. Promote narrowly. Demote when the assumptions stop holding.

That approach is slower than an “enable enforcement” checkbox. It is also easier to audit, easier to falsify and harder to turn into an invisible source of task failure.

Questions developers usually ask

What is Shadow Mode for an AI agent governor?

It is a non-blocking mode where the governor evaluates actions and records recommendations while allowing the agent to continue. That creates evidence about what the governor would have changed before authority is granted.

What is Earned Enforcement?

In MARGINAL, it is the principle that narrow blocking authority must be supported by local evidence and explicit promotion rather than automatically granted at installation.

Should every AI guardrail fail open?

No. Failure semantics depend on the control objective and risk. MARGINAL's fail-open stance is specific to uncertain no-progress efficiency decisions, not a universal security rule.

Try to break the governor before trusting it.

The public “Break MARGINAL” challenge asks for cases where useful verification could look like waste. A counterexample that prevents a bad intervention is more valuable than a flattering benchmark.

Editorial method

This article was developed with AI-assisted drafting and reviewed against MARGINAL's public code, documentation and evidence available on Aug 19, 2026. Product claims are deliberately limited to what those public sources support; unsupported performance claims are excluded.